Multi-modal target quantity statistical method for emergency disaster scene
By employing a multimodal target quantity counting method guided by semantic segmentation and text, the problem of inaccurate target quantity counting in emergency disaster scenarios is solved, achieving high-precision personnel counting and regional rescue support.
Patent Information
- Application Number
- CN202511103352.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies for counting targets in emergency disaster scenarios are easily affected by occlusion, lighting interference, and image damage, leading to inaccurate statistics.
A multimodal target quantity statistical method is adopted. The foreground region mask is extracted by the semantic segmentation branch, and the semantic vector embedding is generated by the text guidance branch. The attention mechanism is used to explicitly fuse the features, input the quantity regression network for prediction, and the model is optimized by an end-to-end training strategy.
It achieves high-precision people counting in complex scenarios, enhances the ability to focus on target areas, reduces misjudgments, is suitable for densely populated and obstructed conditions, and supports regional rescue dispatch and resource allocation.
Smart Images

Figure CN120974202A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision and artificial intelligence technology, and in particular to a method, apparatus, medium and equipment for counting the number of multimodal targets in emergency disaster scenarios. Background Technology
[0002] Emergency rescue is a crucial measure in response to sudden natural disasters, and accurately grasping the distribution of targets within the field of view helps improve the efficiency of emergency rescue. With the development of artificial intelligence technology, existing target density estimation schemes can accurately infer the number of targets within the field of view. Existing technology one proposes a crowd localization algorithm that improves the accuracy of localization counting through a teacher-student knowledge transfer method. However, localization-based methods rely on the regression of detected target boxes and are susceptible to occlusion. Existing technology two proposes a density map-based counting method using segmented branch balanced Gaussian kernel regression. While this method has some robustness to occlusion, it lacks semantic support and struggles to locate the number of local targets. In disaster relief scenarios, image signals are often severely damaged, and the density of people distribution is uneven, necessitating a robust statistical method that integrates semantic, prior, and visual multimodal information. Summary of the Invention
[0003] The main purpose of this application is to provide a method, device, medium and equipment for counting the number of multimodal targets in emergency disaster scenarios, which aims to effectively solve the problem of inaccurate headcount due to factors such as target occlusion, lighting interference and image damage in complex scenarios.
[0004] To achieve the above objectives, this application provides a method for multimodal target quantity statistics in emergency disaster scenarios, comprising: normalizing the original input image and uniformly dividing it into multiple image blocks of the same size according to a preset grid structure, wherein the original input image is an image of an emergency disaster scene; extracting target masks from the foreground region of the original input image through a semantic segmentation branch, and estimating the number of potential individuals in each image block based on the target masks; inputting a preset language template set into a language encoder through a text-guided branch to generate semantic vector embeddings; explicitly fusing the target masks with the backbone visual features of each image block using an attention mechanism to obtain fused features; inputting the fused features and semantic vector embeddings into a quantity regression network to predict the number of targets at the block level, and aggregating the prediction results of all image blocks to obtain a full-image target quantity estimate, thereby obtaining the multimodal target quantity statistics result in the emergency disaster scenario.
[0005] Optionally, the method further includes: employing an end-to-end training strategy, using a comprehensive loss function to jointly optimize and train the semantic segmentation branch, the text guidance branch, the attention mechanism, and the quantitative regression network; wherein the comprehensive loss function includes quantitative regression error, segmentation mask quality error, and text matching error.
[0006] Optionally, the semantic segmentation branch adopts a lightweight deep network structure, extracts the corresponding binary foreground mask from each image block, and uses a connected component analysis algorithm to count the number of connected regions in the binary foreground mask to obtain the number of potential target individuals.
[0007] Optionally, the language template set includes multiple text templates describing the number of targets; the language encoder employs at least one of BERT, GRU, or Transformer architectures.
[0008] Optionally, the explicit fusion process of the attention mechanism includes: extracting the corresponding visual features of each image patch through the backbone network; mapping the target mask to the resolution of the visual features through convolution; weighting the visual features using the query matrix of the attention mechanism and the convolutional features using the key matrix, and then processing them with the softmax function to obtain the attention weights; and multiplying the attention weights, the key matrix of the attention mechanism, and the convolutional features to obtain the explicit fusion result of the attention mechanism.
[0009] Optionally, the quantity regression network adopts a multilayer perceptron structure.
[0010] Optionally, the text matching error is obtained by calculating the similarity between the text semantic vector and the image features.
[0011] Furthermore, to achieve the above objectives, this application also provides a multimodal target quantity counting device for emergency disaster scenarios, comprising: an image block segmentation module for normalizing the original input image and uniformly dividing it into multiple image blocks of the same size according to a preset grid structure, wherein the original input image is an emergency disaster scene image; a foreground extraction module for extracting target masks from the foreground region of the original input image through a semantic segmentation branch, and estimating the number of potential individuals in each image block based on the target masks; a language encoding module for inputting a preset language template set into a language encoder through a text-guided branch to generate semantic vector embeddings; a fusion module for explicitly fusing the target masks with the backbone visual features of each image block using an attention mechanism to obtain fused features; and a prediction module for inputting the fused features and semantic vector embeddings into a quantity regression network to predict the number of targets at the block level, and aggregating the prediction results of all image blocks to obtain a full-image target quantity estimate, thereby obtaining the multimodal target quantity counting results in the emergency disaster scenario.
[0012] To achieve the above objectives, this application also provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the multimodal target quantity statistics method for emergency disaster scenarios provided in the above embodiments.
[0013] To achieve the above objectives, this application also provides an electronic device, which includes: at least one processor, a memory, and an input / output unit; wherein the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the multimodal target quantity statistics method for emergency disaster scenarios provided in any of the foregoing embodiments.
[0014] This application proposes a method, apparatus, medium, and device for multimodal target quantity statistics in emergency disaster scenarios. It normalizes the original input image and divides it into multiple uniformly sized image blocks according to a preset grid structure. The original input image is a picture of an emergency disaster site. An image block partitioning strategy and block-level prediction mechanism are employed to support independent statistics for different regions of the image, facilitating regional rescue scheduling and resource allocation in actual combat. A semantic segmentation branch extracts target masks from the foreground region of the original input image and estimates the potential number of individuals within each image block based on the target masks. A text-guided branch inputs a preset set of language templates into a language encoder to generate semantic vector embeddings. Introducing the text encoder and natural language templates enhances the model's semantic understanding capabilities, enabling it to infer the population distribution even when the image is blurry or lacks local information. An attention mechanism is used to link the target masks with the main visual features of each image block. The method explicitly fuses features to obtain various fused features. It uses the structured mask output from image segmentation and language template embedding to guide feature fusion, which effectively enhances the ability to focus on target regions. It can still achieve high-precision statistics under extreme conditions such as dense crowds and partial occlusion. By explicitly modeling the attention interaction relationship between the mask and the backbone features, the model can still stably identify key regions under conditions such as complex image backgrounds and large changes in lighting, reducing misjudgments. After feature fusion of each fused feature and semantic vector embedding, it is input into the quantity regression network to predict the number of targets at the block level. The prediction results of all image blocks are aggregated to obtain the total number of targets in the image, resulting in multimodal target quantity statistics in emergency disaster scenarios. This method can effectively solve the problem of inaccurate people counting caused by target occlusion, lighting interference, image damage and other factors in complex scenes. It can be smoothly adapted to other target types (such as vehicles, animals, etc.) or deployed in mobile platforms such as drones and vehicle-mounted camera systems, and has broad application prospects. Attached Figure Description
[0015] Figure 1This is a flowchart illustrating an embodiment of the multimodal target quantity counting method for emergency disaster scenarios provided in this application; Figure 2 This is a schematic diagram of an embodiment of the multimodal target quantity statistics method for emergency disaster scenarios provided in this application.
[0016] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0017] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0018] This application proposes a multimodal target quantity counting method suitable for emergency disaster environments. Combining semantic segmentation, a language-guided mechanism, and an attention-enhanced backbone regression branch, it effectively addresses the inaccurate number counting problem caused by target occlusion, lighting interference, and image damage in complex scenes. This method overcomes the limitations of traditional detection-classification structures by integrating region mask priors, language knowledge guidance, and deep feature regression into a unified framework, improving the perception of local and overall target distribution. The core innovations of this method are: firstly, the introduction of a language template guidance mechanism, which establishes semantic assumptions about target quantity through pre-set text sentences, thereby assisting the network in judging target density in regions when visual information is limited; secondly, the design of an explicit attention interaction structure, which fuses mask information obtained from the semantic segmentation branch with backbone network features, enabling the model to focus on the actual target location during region representation and effectively suppressing background interference. Furthermore, this method outputs quantity predictions on an image patch basis, achieving partitioned evaluation and supporting full-image aggregation, demonstrating strong practicality and interpretability. Furthermore, this application employs an end-to-end joint training mechanism, integrating regression error, segmentation quality, and text matching accuracy through multi-task loss to further enhance the overall robustness of the model. It can be widely applied to personnel search and rescue statistics in emergency scenarios such as earthquakes, fires, and floods, providing an intelligent perception foundation for subsequent emergency response.
[0019] The technical solution adopted by this application to solve its technical problem is: a multimodal target quantity statistics technology for emergency rescue, characterized by the following steps: Reference Figure 1 The first embodiment of this application provides a method for counting the number of multimodal targets in emergency disaster scenarios. Figure 2 This is a schematic diagram of a multimodal target quantity counting method for emergency disaster scenarios. This method can be executed by a terminal or server processor. Multimodal target quantity counting methods for emergency disaster scenarios may include: S10. Normalize the original input image and divide it into multiple image blocks of the same size according to a preset grid structure. The original input image is an emergency disaster scene image. This application first normalizes the original input image to unify the input dimension and improve the model's adaptability to scale changes. To support population estimation in local areas and enhance the system's spatial awareness, the input disaster image is then... Divide the grid evenly according to the preset grid structure. A uniformly sized image patch:
[0020] Each image block This serves as the processing unit for subsequent branch modules. This division method can improve the system's ability to model local crowd density while maintaining the overall semantics of the image.
[0021] For example, video images acquired in emergency scenarios typically have a large field of view, and the targets exhibit scale variations. Therefore, image standardization is necessary to unify the input image to a fixed resolution (e.g., 512×512), and then divide it into... Each image patch (e.g., 8×8) serves as the smallest unit for independent processing of subsequent branches, thus supporting population estimation at the region level.
[0022] Clearly, this segmentation method can improve the system's ability to model local crowd density while maintaining the overall semantics of the image.
[0023] S20. Extract the target mask of the foreground region from the original input image through the semantic segmentation branch, and estimate the number of potential individuals in each image block based on the target mask; In one embodiment of this application, the semantic segmentation branch adopts a lightweight deep network structure, extracts the corresponding binary foreground mask from each image block, and uses a connected component analysis algorithm to count the number of connected regions in the binary foreground mask to obtain the number of potential target individuals.
[0024] Specifically, after image patch segmentation, a dedicated semantic segmentation branch is constructed to extract mask regions of the foreground regions (i.e., targets such as heads and bodies) in the image. This segmentation module can employ a lightweight deep network, such as U-Net or SegNet, to extract the mask of the foreground regions (i.e., the detected objects) from the original image I.
[0025] in, Represents a segmented network. It is a binary mask image. Represents image blocks The corresponding local region. Through analysis... Number of connected regions This allows us to estimate the number of potential individuals within the block:
[0026] in, It is a non-linear mapping function that can be used for empirical modeling or regression learning based on specific tasks.
[0027] For example, the processor can use a lightweight segmentation network such as U-Net to output a binary foreground mask from the original image.
[0028] In the foreground mask, a value of "1" corresponds to a location where a human head might be present, while other areas are the background.
[0029] Next, for each image patch, a local region is cropped from the global mask. The number of connected regions was counted using a connected component analysis algorithm. Based on this, a rough prior number of people can be obtained:
[0030] in It is a non-linear mapping function that can be used for empirical modeling or regression learning based on specific tasks.
[0031] S30. Input the preset language template set into the language encoder through the text-guided branch to generate semantic vector embedding; In one embodiment of this application, the language template set includes multiple text templates describing the target quantity; the language encoder employs at least one of BERT, GRU, or Transformer architectures.
[0032] To further enhance the model's understanding of the concept of people, a language template mechanism is introduced, transforming the target quantity statistics task into an image-text matching problem. A template set is defined. The format is "There are c people in this area", where Each template is processed by a language encoder. After processing, semantic vector embeddings are generated:
[0033] Where d represents the language embedding dimension, and the language encoder can be a BERT, GRU, or Transformer architecture, generating... It represents the semantic intent under this quantity assumption.
[0034] For example, the processor can input a pre-defined number of people prompt text (such as "There are c people in this area") into the language encoder model. The language encoder can use a pre-trained model BERT or a simplified GRU network, which outputs a... Fixed-dimensional semantic vectors .
[0035] S40. Explicitly fuse the target mask with the backbone visual features of each image patch using an attention mechanism to obtain each fused feature; The explicit fusion process of the attention mechanism includes: The corresponding visual features of each image patch are extracted through the backbone network; Map the target mask to the resolution of visual features through convolution; After weighting the visual features with the query matrix and the convolutional features with the key matrix, the attention weights are obtained by applying the softmax function. The explicit fusion result of the attention mechanism is obtained by multiplying the attention weights, the key-value matrix of the attention mechanism, and the convolutional features.
[0036] Specifically, after acquiring the segmentation mask and the backbone visual features of the image patch, the processor explicitly fuses them using an attention mechanism. The backbone network then processes the image patch... Extracting visual features The mask region is processed by convolution. After mapping to the feature resolution, the attention weights are constructed:
[0037] in, This is a learnable matrix, representing the mapping operation between the query and the key. Then, the fused features are calculated:
[0038] Final output It can be used for subsequent prediction tasks.
[0039] S50. After fusing the fused features with the semantic vector embeddings, input the results into the quantity regression network to predict the number of targets at the block level. Then, aggregate the prediction results of all image blocks to obtain the total number of targets in the entire image, thus obtaining the multimodal target quantity statistics in the emergency disaster scenario. The quantity regression network employs a multilayer perceptron structure.
[0040] In one embodiment of this application, the multimodal target quantity counting method for emergency disaster scenarios further includes: An end-to-end training strategy was adopted, and a comprehensive loss function was used to jointly optimize and train the semantic segmentation branch, the text guidance branch, the attention mechanism, and the quantitative regression network. The comprehensive loss function includes quantitative regression error, segmentation mask quality error, and text matching error.
[0041] Specifically, to simultaneously optimize the performance of segmentation, image-text fusion, and quantity prediction, this application designs a comprehensive loss function for end-to-end joint training. The comprehensive loss function includes: quantity regression error. Segmentation mask quality error Text matching error :
[0042] The meanings of each sub-loss are as follows: Represents the supervised regression error, where Label the actual number of people; The segmentation loss function, usually IoU loss or Dice loss, measures the degree of regional overlap between the predicted mask and the ground truth. For text matching loss, such as image-text embedding distance or multimodal contrast loss, it is used to guide... and Semantic consistency.
[0043] For example, the processor will fuse the image features With corresponding text semantic vector The data is then spliced or fused with attention and fed into the quantity regression network module R. The network can employ a multilayer perceptron structure to predict the quantity of block-level targets.
[0044] in Indicates the first The number of individuals predicted within an image patch. This is an estimate of the total population of the entire map.
[0045] The effectiveness of this application can be further illustrated by the following simulation experiments.
[0046] The simulation hardware in this application includes an Intel(R) Xeon(R) CPU E5-2680 v4 2.40GHz, 128G of memory, and a Linux operating system, and the simulation is performed using the deep learning framework PyTorch.
[0047] The data used in the simulation of this application are the publicly available datasets SHHA and SHHB, specifically derived from publicly available population flow data in Shanghai streets in the prior art.
[0048] This experiment uses two commonly used evaluation metrics in target counting tasks: Mean Absolute Error (MAE) and Mean Squared Error (MSE). MAE represents the difference between the predicted number of people and the labeled number of people, and its definition is as follows:
[0049] in, This indicates the predicted population size. This represents the number of real people labeled. MSE is used to measure the dispersion of MAE across all test samples, and its definition is as follows:
[0050] Table 1 compares the experimental results of the proposed method and existing method 1 on the SHHA dataset, demonstrating the improvement in target count brought about by the introduction of semantic information in existing algorithms. Table 2 compares the experimental results of the proposed method and existing method 2 on the SHHB dataset, further proving the advancement and effectiveness of the segmentation mask-guided multimodal target count method proposed in this application.
[0051] Table 1 Comparison of Experimental Results
[0052] Table 2 Comparison of Experimental Results II
[0053] The second embodiment of this application provides a multimodal target quantity counting device for emergency disaster scenarios. The device may include: an image block segmentation module for normalizing the original input image and uniformly dividing it into multiple image blocks of uniform size according to a preset grid structure, wherein the original input image is an image of an emergency disaster scene; a foreground extraction module for extracting target masks from the foreground region of the original input image through a semantic segmentation branch, and estimating the number of potential individuals within each image block based on the target masks; a language encoding module for inputting a preset language template set into a language encoder through a text-guided branch to generate semantic vector embeddings; a fusion module for explicitly fusing the target masks with the main visual features of each image block using an attention mechanism to obtain fused features; and a prediction module for inputting the fused features and semantic vector embeddings into a quantity regression network to predict the number of targets at the block level, and aggregating the prediction results of all image blocks to obtain a full-image target quantity estimate, thus obtaining the multimodal target quantity counting result in the emergency disaster scenario.
[0054] The third embodiment of this application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the multimodal target quantity statistics method for emergency disaster scenarios described in any of the preceding claims.
[0055] The fourth embodiment of this application provides an electronic device, characterized in that the electronic device includes: at least one processor, a memory, and an input / output unit; wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the aforementioned multimodal target quantity statistics method for emergency disaster scenarios.
[0056] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for counting the number of multimodal targets in emergency disaster scenarios, characterized in that, include: The original input image is normalized and evenly divided into multiple image blocks of the same size according to a preset grid structure. The original input image is a picture of an emergency disaster scene. The semantic segmentation branch extracts the target mask of the foreground region from the original input image, and estimates the number of potential individuals in each image block based on the target mask; The pre-defined language template set is input into the language encoder through the text-guided branch to generate semantic vector embeddings; An attention mechanism is used to explicitly fuse the target mask with the backbone visual features of each image patch to obtain the fused features. After feature fusion by embedding each fused feature with a semantic vector, the input is fed into a number regression network to predict the number of targets at the block level. The prediction results of all image blocks are then aggregated to obtain the total number of targets in the image, thus obtaining the statistical results of the number of multimodal targets in emergency disaster scenarios.
2. The multimodal target quantity statistics method for emergency disaster scenarios as described in claim 1, characterized in that, The method further includes: An end-to-end training strategy was adopted, and a comprehensive loss function was used to jointly optimize and train the semantic segmentation branch, the text guidance branch, the attention mechanism, and the quantitative regression network. The comprehensive loss function includes quantitative regression error, segmentation mask quality error, and text matching error.
3. The multimodal target quantity statistics method for emergency disaster scenarios as described in claim 1, characterized in that, The semantic segmentation branch utilizes a lightweight deep network structure to extract the corresponding binary foreground mask from each image block, and uses a connected component analysis algorithm to count the number of connected regions in the binary foreground mask to obtain the number of potential target individuals.
4. The multimodal target quantity statistics method for emergency disaster scenarios as described in claim 1, characterized in that, The language template set includes multiple text templates describing the target quantity; the language encoder employs at least one of BERT, GRU, or Transformer architectures.
5. The method for counting the number of multimodal targets in emergency disaster scenarios as described in claim 1, characterized in that, The explicit fusion process of the attention mechanism includes: The corresponding visual features of each image patch are extracted through the backbone network; Map the target mask to the resolution of visual features through convolution; The attention weights are obtained by weighting the visual features using the query matrix and the convolutional features using the key matrix, and then processing them with the softmax function. The explicit fusion result of the attention mechanism is obtained by multiplying the attention weights, the key-value matrix of the attention mechanism, and the convolutional features.
6. The method for counting the number of multimodal targets in emergency disaster scenarios as described in claim 1, characterized in that, The quantitative regression network employs a multilayer perceptron structure.
7. The method for counting the number of multimodal targets in emergency disaster scenarios as described in claim 2, characterized in that, The text matching error is obtained by calculating the similarity between the text semantic vector and the image features.
8. A multimodal target quantity counting device for emergency disaster scenarios, characterized in that, include: The image block segmentation module is used to normalize the original input image and divide it into multiple image blocks of the same size according to a preset grid structure. The original input image is an emergency disaster scene image. The foreground extraction module is used to extract the target mask of the foreground region from the original input image through the semantic segmentation branch, and estimate the number of potential individuals in each image block based on the target mask; The language encoding module is used to input a pre-defined set of language templates into the language encoder through the text-guided branch to generate semantic vector embeddings. The fusion module is used to explicitly fuse the target mask with the backbone visual features of each image patch using an attention mechanism to obtain fused features. The prediction module is used to input the fused features and semantic vector embeddings into the quantity regression network to predict the number of targets at the block level, and aggregate the prediction results of all image blocks to obtain the total number of targets in the whole image, thus obtaining the statistical results of the number of multimodal targets in emergency disaster scenarios.
9. A computer-readable storage medium, characterized in that, It includes instructions that, when run on a computer, cause the computer to execute the multimodal target quantity statistics method for emergency disaster scenarios as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, The electronic device includes: At least one processor, memory, and input / output unit; The memory is used to store computer programs, and the processor is used to call the computer programs stored in the memory to execute the multimodal target quantity statistics method for emergency disaster scenarios according to any one of claims 1 to 7.