Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

64 results about "Crowd counting" patented technology

Crowd counting or crowd estimating is a technique used to count or estimate the number of people in a crowd. The most direct method is to actually count each person in the crowd, for example turnstiles are often used to precisely count the number of people entering an event.

Video crowd counting method based on cascaded cross-domain feature interaction network

The invention discloses a video crowd counting method based on a cascaded cross-domain feature interaction network. The method comprises the following steps: carrying out data enhancement processing of random cutting and horizontal flipping on a current frame and front and back frames of the current frame; and constructing a cross-domain feature interaction network composed of a spatial domain branch and a frequency domain branch. The frequency domain branch extracts frequency domain feature output of different stages through a high and low frequency signal aggregation module and a feature encoder based on adjacent frames; the spatial domain branch is based on a single-frame image, and static spatial semantic features are extracted through a feature encoder. Cascade fusion is carried out on the double-branch features on multiple scales, two-way channel cross attention is utilized to reconstruct time sequence correlation frequency domain features of a current frame, and fusion and reconstruction of the two domain features are achieved through a cross-domain feature mutual modulation module. And after the reconstructed double-branch features are processed by the fusion network, outputting a crowd density map of the current frame by a density regression head. And after training is completed, storing the optimal model for video crowd counting. According to the invention, through cross-domain feature cascade and bidirectional time sequence modeling, the accuracy and robustness of crowd counting in a video scene are effectively improved.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Lightweight crowd counting and positioning method

The invention discloses a lightweight crowd counting and positioning method, and the method comprises the steps: building a brand-new lightweight crowd counting and positioning network, designing a grouping feature pairing interaction module, carrying out the grouping of feature maps, splicing adjacent groups, learning the context relation between feature groups, generating a dynamic attention weight to enhance the expression of key features, and carrying out the recognition of the key features. Therefore, the recognition capability of the target in the complex shielding scene is improved. A multi-stage training strategy is adopted, backbones, positioning branches, segmentation branches and lightweight adapter modules are sequentially and independently trained, and stable convergence of the multi-task capability of the model is ensured; according to the method, excellent performance is achieved on the public crowd counting data set, counting and positioning precision is improved, low calculation overhead is kept, and an efficient and accurate solution is provided for crowd analysis tasks in a resource limited scene.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY

Crowd counting method and system based on WiFi and video modal cross-level attention

The invention relates to the technical field of crowd counting, in particular to a crowd counting method and system based on WiFi and video modal cross-level attention, and the method comprises the steps: constructing a WiFi density map at a WiFi sensing side, and converting an irregular detection record into a fixed-size image representation; on a video sensing side, marking a region of interest for video frames collected by cameras with different visual angles, and cutting the region of interest to serve as a video side image; respectively carrying out feature coding on the WiFi density map and the video side image by adopting a convolutional neural network and self-attention combined mode; gradually aligning the WiFi modal feature embedded representation and the video modal feature embedded representation through multi-layer stacked cross-modal attention to obtain a cross-modal fusion feature; and inputting the cross-modal fusion features into a lightweight multilayer perceptron, and outputting crowd count. According to the invention, through hierarchical alignment and fusion of WiFi signals and video features, accurate estimation of the number of crowds in a large-scale complex scene can be realized.
Owner:INNER MONGOLIA ZHIXING HUILIAN TECHNOLOGY CO LTD

RGB-T counting method and device based on semantic perception complementary feature mining

The invention discloses an RGB-T counting method and device based on semantic perception complementary feature mining. The method comprises the steps that firstly, images in a training data set are preprocessed; then constructing an RGB-T counting network based on semantic perception complementary feature mining, wherein the RGB-T counting network comprises a feature extractor based on Transform, a mode feature adaptive method from coarse to fine and a semantic perception agent module; the feature extractor based on the Transform is used for extracting high-level semantic representations of two different modal images; the mode feature adaptive method from coarse to fine is used for filtering noise features and mining fine complementary features; and the semantic perception agent module is used for enhancing semantic consistency perception in network output and the feature extractor, and predicting a density map and a counting result through a density regression layer. According to the method, the RGB-T crowd counting task can be efficiently completed, and the counting result is better than that in the prior art.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

A multi-modal crowd density prediction method based on time convolution network

The application discloses a kind of multi-modal crowd density prediction methods based on time convolution network, comprising the following steps: video information module obtains image from monitoring camera, the image obtained is carried out crowd counting by crowd counting model, obtains the crowd density of each camera every time, and is organized into time series data;The time series data of crowd density is respectively extracted into the current time hidden vector for each sub-region by time convolution network;Itinerary planning module extracts corresponding itinerary information from itinerary planning table for multi-modal prediction;The features of fusion module and video information module are fused, and the crowd density prediction value of each sub-region in the future is obtained.This application can improve the prediction accuracy of the model by fusing multi-modal information, and can give a smooth prediction result when facing low SNR data, and can also respond in time when the crowd density changes suddenly.
Owner:SOUTH CHINA UNIV OF TECH

Domain-oriented adaptive crowd counting energy-driven active learning method

The invention belongs to the technical field of crowd counting, and particularly relates to a domain-oriented adaptive crowd counting energy-driven active learning method, which comprises the following steps of: extracting multi-scale visual features from an input image, and training HRNet to obtain an optimal source domain model; source domain and target domain samples are mapped to a unified energy space through an energy function; performing data enhancement on the target domain sample, and calculating uncertainty and a predicted people number mean value under various data enhancement modes; calculating sample energy of the target domain; screening the target domain samples twice by adopting an active learning strategy, and labeling the screened samples; designing energy alignment loss; and performing fine tuning on the optimal source domain model, and obtaining a final crowd counting result by using the fine-tuned model. According to the method, data distribution of the source domain and the target domain is effectively aligned through the energy model, sample labeling is carried out in combination with an active learning strategy, and the cross-domain crowd counting performance is improved.
Owner:SHANDONG UNIV OF TECH

Crowd counting method and device, terminal equipment and storage medium

The application discloses a crowd counting method and device, a terminal equipment and a storage medium. The method comprises the following steps: acquiring a target picture; performing prediction on the target picture by using a pre-trained crowd counting model to obtain a density map prediction value and a head number prediction value; and performing weighted calculation on the density map prediction value and the head number prediction value to obtain a crowd quantity count value. The head number prediction value is used as an auxiliary result to be weighted with the density map prediction value, so that the accuracy of the obtained crowd quantity count value is improved, and the accuracy of crowd counting is improved.
Owner:CHINA MOBILE GROUP JIANGSU +1

Dense crowd counting method based on improved yolov11 model

The invention discloses a dense crowd counting method based on an improved yolov11 model. The dense crowd counting method comprises the following steps: improving a feature extraction network of a yolov11n model; replacing an original CBS and C3K2 module of the yolov11n model with an encoder module of a U-NET V2 network; the feature fusion network of the yolov11n model is improved; an AIFI module is added behind the highest layer of the feature fusion network; a detection head of the yolov11n model is improved, and a DIOU loss function is adopted to replace an original loss function; and inputting the preprocessed image tensor into the improved yolov11n model for training, and storing the trained model for dense target counting detection. According to the method, the calculation speed is ensured as much as possible, and the precision of crowd number detection when severe shielding exists between targets in a dense scene is improved.
Owner:BEIJING UNIV OF TECH

Traffic hub crowd counting method and system based on multi-scale cross-region graph convolution driving

ActiveCN121392731BCrowd countingData set
The present application relates to the field of transportation information engineering, and more particularly to a traffic hub crowd counting method and system based on multi-scale cross-region graph convolution driving, wherein the method comprises: constructing a crowd counting sequence; constructing three time series density map identification evaluation standards of spatial consistency, time sequence continuity and scale adaptability; proposing a multi-scale cross-region graph convolution network model to train and test the traffic hub crowd counting; and performing instance verification of the traffic hub crowd counting. The proposed model is applied to a series of empirical data sets for training and testing, which proves that it can achieve better demand prediction performance than the baseline model.
Owner:CHINA FIRST HIGHWAY ENGINEERING CO LTD +2

A crowd counting method based on perspective self-adaption in complex scene

The application discloses a crowd counting method in a complex scene based on a visual angle self-adaption, and the counting method comprises the following steps: a NOOMP framework is fitted to a natural world through a few-shot learning method; the NOOMP framework is trained through meta-learning to adapt a multi-head parallel network to a main body of the NOOMP framework; the multi-head parallel network is used for estimating a density map of the crowd; and the multi-head parallel network is trained in multiple different scenes, and sub-losses are summarized. The application proposes a new marking method, an absolute geometric Gaussian generation method, and the method can obtain better precision by only adding a point to each person in an image.
Owner:NANCHANG UNIV

A personnel off-duty detection method based on Hungarian algorithm and P2PNet

The present application relates to the technical field of post safety management, and specifically relates to a personnel off-duty detection method based on a Hungarian algorithm and a P2PNet, which comprises the following steps: step 1, acquiring image information, installing a camera, adjusting the irradiation direction of the camera so that it includes all posts in the monitored area, and collecting image information of on-site personnel under different conditions; step 2, data labeling and model training, labeling the original data set, and training a head center point detection model based on the P2PNet; the present application uses artificial intelligence technology to detect personnel off-duty, combines the crowd counting P2PNet algorithm and the Hungarian matching algorithm, and is deployed on the Cambrian MLU370-S4 intelligent acceleration card, thereby realizing real-time automatic detection of personnel off-duty. The method has high robustness and reliability for various scenes, and the intelligent acceleration card ensures the timeliness of the detection, reduces a large amount of labor cost, and ensures the safety of production operations.
Owner:GUONENG JIANGXI NEW ENERGY IND CO LTD

A weakly supervised crowd counting method based on multi-scale dynamic graph convolution

A weakly supervised crowd counting method based on a multi-scale dynamic graph convolution network belongs to crowd counting in the fields of public security, city planning and traffic scheduling. Due to the complexity and diversity of traffic scenes, it is very difficult to perform point-level labeling on a large number of crowds, and a large amount of manpower is required. Weakly supervised crowd counting is more suitable for these scenes because they only require counting-level annotations. Existing weakly supervised crowd counting ignores the non-uniformity of cross-distance crowd density distribution and multi-scale crowd head, and cannot obtain similar accurate counting results as the fully supervised crowd counting method. The present application proposes a multi-level regional dynamic graph convolution module to extract the internal relationship between different crowd regions, so as to learn dynamic regional scores and further optimize regional feature representation, and a coarse-grained multi-level feature fusion module is designed to extract multi-scale crowd head information. The present application has high regression accuracy and end-to-end crowd counting capability.
Owner:BEIJING UNIV OF TECH

A crowd counting method and device, electronic equipment and storage medium

The present disclosure relates to a crowd counting method and device, an electronic device and a storage medium. The method comprises: obtaining a crowd image; obtaining a first number of people corresponding to the crowd image and a first crowd density distribution map corresponding to the crowd image based on head key point positioning on the crowd image; obtaining a second crowd density distribution map corresponding to the crowd image based on crowd density detection on the crowd image; selecting a target crowd density distribution map corresponding to the crowd image from the first crowd density distribution map and the second crowd density distribution map based on the first number of people and a first preset number threshold; and determining a crowd counting result of the crowd image based on the target crowd density distribution map. The embodiments of the present disclosure can improve the accuracy of crowd counting in various scenarios.
Owner:SHANGHAI SENSETIME INTELLIGENT TECH CO LTD

A multi-scale alignment fusion-based multi-modal crowd counting method

This invention provides a multimodal crowd counting method based on multi-scale alignment and fusion. First, a multimodal crowd scene dataset is acquired and divided into training, validation, and test sets. Preprocessed RGB images and multimodal auxiliary images are input into a VGG16 backbone network to extract multi-stage high-level feature maps. These are then sequentially fed into a density-sharing local contrastive learning module, a local feature fusion module, and an adaptive Mamba context-aware fusion module to achieve cross-modal fine-grained alignment, local feature fusion, and global context modeling. Subsequently, the global fused features are input into a dynamically upsampled multi-scale feature decoder to generate a high-resolution crowd density map. Supervised training is performed using a composite loss function, and the optimal model is saved for testing. This invention adopts an "align-then-fuse" architecture, effectively mitigating cross-modal heterogeneity and density fluctuation problems through multi-module collaborative design, significantly improving the accuracy and robustness of crowd counting in complex scenes.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A video crowd counting method based on time-series interaction and global correlation network

This invention discloses a video crowd counting method based on temporal interaction and a global association network, belonging to the field of crowd counting technology. It involves acquiring a continuous video sequence of a preset length containing pedestrians, forming a sample set with T consecutive video frames and the ground truth density map corresponding to each of the T consecutive video frames. Based on the sample set, a training model including an encoding module, a dual-branch feature fusion module, a channel-guided cross-branch feature fusion module, and a feature integration module is trained to obtain a temporal interaction and global association network used to generate the predicted density map of T consecutive frames. The T consecutive video frames to be estimated are input into the trained temporal interaction and global association network to obtain the predicted density map of T consecutive frames. The crowd estimation result is generated by summing the results frame by frame. This invention effectively improves the accuracy and robustness of crowd counting in video scenes through dual-branch parallel processing and cross-dimensional feature fusion.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Unmanned aerial vehicle aerial image dense crowd counting and grade classification method based on multilayer CNN

The invention discloses an unmanned aerial vehicle aerial image dense crowd counting and grade classification method based on a multilayer CNN, and the method comprises the following steps: carrying out the standardization preprocessing of an input unmanned aerial vehicle aerial image I; multi-scale features of the image are extracted through multi-layer convolution and pooling operation, and a high-level semantic feature map F is obtained; a context aggregation module is utilized, adaptive average pooling of different scales is adopted to capture global to local context information, and a feature map F'with attention weight is generated through convolution fusion; a rear-end decoder recovers spatial details by using cavity convolution, a single-channel density map D is output through convolution operation, and a crowd counting result is obtained through integration; and finally, based on the density map D, dividing four density intervals through a threshold value: analyzing and marking a region above medium density by using 8 connected domains, and outputting a visualization result on the image I. According to the invention, grading and counting of crowd density in an unmanned aerial vehicle aerial photographing scene can be accurately realized, and applications such as public safety monitoring and large-scale activity management can be effectively supported.
Owner:SHENYANG AEROSPACE UNIVERSITY

A new people counting method

The present application provides a new crowd counting method. First, the first 16 layers of VGG19 are used as the backbone network to extract shallow features, and then a double-branch structure is used in the feature extraction module. Branch 1 uses a pyramid structure with a fusion self-attention mechanism, and the feature map generated by the pyramid structure is sent to a transition residual block to generate the feature map of branch 1. Branch 2 uses a double-channel attention module, and the feature maps obtained by branches 1 and 2 are sent to a transition residual block for splicing and fusion. The fused feature map is sent to a transition module to generate the final feature map. Finally, the final feature map is sent to a 1x1 convolution to generate a density map. During model training, the present application uses a joint loss function to minimize the influence of outliers on the entire model. Branch 1 of the present application can accurately locate targets of different scales and depict the spatial dependency between any two positions in the feature map. Branch 2 of the present application can focus on important features in the crowd, thus achieving excellent crowd counting performance.
Owner:SOUTHWEAT UNIV OF SCI & TECH

Method for constructing multi-dimensional dynamic perception and progressive focusing network model for crowd counting

The application provides a construction method of a multi-dimensional dynamic perception and progressive focusing network model for crowd counting, and relates to the technical field of computer vision, and comprises: a front-end sub-network used for shallow image feature extraction on a preprocessed input image to obtain image shallow features; a main body sub-network used for multi-branch feature interaction based on the image shallow features to generate deep features and cross-layer features, cross attention processing and multi-scale feature fusion on the deep features and the cross-layer features, cross-dimensional attention modeling on the fused features, and layer normalization output of depth features; and a rear-end sub-network used for high-dimensional feature construction and spatial resolution adjustment on the depth features to obtain intermediate features, feature optimization on the intermediate features, attention fusion and density decoding, and output of a density map and a target attention map. The network model can realize crowd counting in a wide-area scene.
Owner:SOUTH CENTRAL UNIVERSITY FOR NATIONALITIES

Dense crowd counting method combining high-resolution CNN and lightweight transformer

The application provides a dense crowd counting method combining a high-resolution CNN and a lightweight Transformer, and comprises the following steps: the scale size of a human head in a crowd image is calculated by using a fixed Gaussian kernel method to generate a supervised density map for network training; a crowd counting network based on a high-resolution feature extraction network HRNet and a lightweight Transformer is constructed; data augmentation is performed on a crowd data set, the constructed counting network is trained by using a training set, and an optimal model is screened and saved; the optimal network model obtained is tested by using a test set, and the final counting result of the picture crowd is obtained by accumulating and summing the pixel values of the network predicted density map. The application can not only maintain high-resolution output of crowd features, but also can fuse multi-scale information, improve the robustness of crowd counting, and significantly improve the convergence speed and generalization performance of the model.
Owner:SICHUAN UNIV

Crowd counting method, device, equipment and medium

The invention discloses a crowd counting method, apparatus and device, and a medium. The crowd counting method comprises the steps of obtaining a crowd input image; performing feature extraction on the crowd input image by using a first convolution kernel based on a first void rate by using a first void convolution hierarchy to obtain a first feature map; performing down-sampling feature extraction on the first feature map by using a second convolution kernel based on a second void rate by using a second void convolution hierarchy to obtain a second feature map; performing feature fusion processing based on the crowd input image, the first feature map and the second feature map to obtain a dilated convolution fusion feature; performing asymmetric convolution fusion processing on the second feature map and the dilated convolution fusion feature to obtain a target fusion feature; and performing mapping regression processing on the target fusion features to obtain a crowd counting estimation result.
Owner:BEIJING JIZHI DIGITAL TECH CO LTD

A region complement-based population counting method and related device

The application discloses a crowd counting method based on regional complementary aggregation and a related device. The method comprises the following steps: acquiring a dense crowd image; inputting the dense crowd image into a regional complementary aggregation network model; determining a fine density map through the regional complementary aggregation network model; and determining a crowd counting result corresponding to the dense crowd image based on the fine density map. The application generates a coarse density map through bidirectional iterative fusion and weighted complementary cascading of a complementary iterative aggregation module in the regional complementary aggregation network model, determines an attention map through multi-depth supervision of a region positioning module in the regional complementary aggregation network model, and then determines a fine density map based on the attention and the coarse density map. Through multi-scale fusion of the complementary iterative aggregation module and multi-depth supervision of the region positioning module, the image quality of the fine density map is effectively improved, so that the accuracy of the crowd counting result determined based on the fine density map can be improved.
Owner:PENG CHENG LAB

Crowd counting method based on multi-feature fusion VMama

The invention relates to the technical field of computer vision, in particular to a crowd counting method based on multi-feature fusion VMama, and the method comprises the steps: obtaining a to-be-counted crowd image data set, carrying out the resolution normalization processing of an image, and constructing an end-to-end crowd counting model with VMama as a backbone network, extracting four-stage multi-scale features of the image through a VMama backbone network; performing fusion operation through a multi-feature fusion module to generate multi-scale fusion features; inputting the fusion features into an integration attention module, and outputting target enhancement features through feature weight distribution of channel attention and spatial positioning enhancement of coordinate attention; inputting the target enhanced features into an ellipse constraint deformable convolution module, and executing expansion convolution processing on output features of the ellipse constraint deformable convolution module to generate a crowd density map; and the pixel values of the density map are summed to obtain a final crowd counting result, so that the crowd counting effect is improved.
Owner:HENAN UNIVERSITY

Multi-element example unified crowd counting method and system based on visual language prompt

The invention discloses a multi-element example unified crowd counting method and system based on visual language prompt, and belongs to the technical field of crowd counting. The method comprises the following steps: inputting a target image into an image encoder to capture local detail features and global context information to obtain image embedding, carrying out position encoding on a point input prompt and a box input prompt through a position encoder to obtain position embedding, and carrying out text encoding on a text input prompt through a text encoder to obtain text embedding, the position embedding and the text embedding are input into a prompt embedding module for merging to obtain prompt embedding, the image embedding and the prompt embedding are input into a task adaptive decoder, and the image embedding and the prompt embedding are fused through a bidirectional attention feature fusion module to obtain shared feature representation; and inputting the shared feature representation into a corresponding target prediction head module in a task decoder based on the input prompt to obtain a crowd quantity prediction result. The method improves the accuracy of crowd counting.
Owner:GUANGDONG HUST IND TECH RES INST +1

A multi-scale based feature extraction crowd counting method and system

The application is suitable for the technical field of crowd counting, and provides a crowd counting method and system based on multi-scale feature extraction, which comprises the following steps: extracting primary features of an input image through a VGG-16 backbone network; inputting the primary features into MSGM, wherein the MSGM comprises a SAFMN and an LSGA; performing multi-scale feature fusion and modulation on the input primary features by using the SAFMN, inputting the features output by the SAFMN into the LSGA, modeling spatial correlation through a Gaussian position matrix, and generating an optimized feature map; predicting a crowd density distribution based on the optimized feature map, and calculating a crowd quantity; and through the cross-scale feature aggregation of the SAFMN module and the Gaussian attention mechanism of the LSGA module, the application effectively solves the target scale difference problem in the monitoring image, and improves the feature extraction capability for multi-scale head targets.
Owner:MINNAN INST OF SCI & TECH

Crowd counting system and method based on cross-modal feature registration and ghosting suppression

The invention provides a crowd counting system and method based on cross-modal feature registration and ghosting suppression, and relates to the technical field of computer vision. The system comprises a visible light and infrared light feature extraction module, three cross-modal feature registration modules, four cross-modal ghosting suppression and fusion modules and a crowd density map estimation module. The visible light and infrared light feature extraction module extracts multi-level features from visible light and infrared light images of the same scene. The cross-modal feature registration module performs spatial registration on the visible light features based on the infrared light features so as to eliminate dislocation between modals; the ghosting suppression and fusion module carries out redundancy suppression and attention fusion on registered or unregistered features at all feature levels to generate multi-level fusion features; and the density map estimation module carries out crowd density estimation according to the multi-level fusion features to obtain the estimated value of the number of people in the scene to be counted. According to the method, the precision and robustness of crowd counting are remarkably improved.
Owner:YANSHAN UNIV

Video crowd counting method, device, terminal equipment and storage medium

The application discloses a video crowd counting method, device and equipment and a storage medium, which comprises the following steps: inputting image features of a to-be-counted image sequence corresponding to a to-be-counted image into a decoder of a target deep neural network model; extracting first spatial features in the to-be-counted image features through a local spatial self-attention module; extracting first time features in the to-be-counted image features through a global time self-attention module; generating a first crowd density map based on the first spatial features and the first time features through the target deep neural network model, and determining a target crowd density map corresponding to the to-be-counted image sequence based on the first crowd density map; and adding the target crowd density map corresponding to the to-be-counted image sequence pixel by pixel through the target deep neural network model to obtain a crowd counting result corresponding to the to-be-counted image sequence. The application realizes the space-time correlation between image sequences in the crowd counting algorithm and improves the counting accuracy of the algorithm.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL

A method for constructing a multi-modal crowd counting model

ActiveCN115359428BCrowd countingEngineering
The application discloses a kind of multi-modal crowd counting model construction methods, comprising: extracting multi-modal feature from multi-modal source signal;Set learnable counting feature;Cascade multi-modal feature and counting feature, form the fusion feature of counting guidance;Through multi-head self-attention block, the fusion feature of counting guidance is enhanced, and enhanced feature is formed;Split enhanced feature, form enhanced multi-modal feature and enhanced counting feature;Enhanced multi-modal feature channel cascade is used Prediction head carries out the prediction of density map;Using multilayer perception, enhanced counting feature is reduced channel, and forms count value;Using density map true value supervises density map, using count value supervises the count value of density map statistics, using count value supervises;Through training set training forms multi-modal crowd counting model.The model constructed by the application can improve crowd counting precision by the guidance of counting information, multi-modal fusion is implemented by multi-head self-attention.
Owner:ANHUI UNIV

An unsupervised population counting method, device and storage medium

This invention discloses an unsupervised crowd counting method, apparatus, and storage medium. The method includes cropping a first input image into image blocks and obtaining coarse-grained text for each image block; inputting the image blocks into a first image encoder and the coarse-grained text into a first text encoder to generate a first similarity matrix; filtering image blocks of a first target category based on the first similarity matrix and a first discriminative category similarity; obtaining fine-grained text for image blocks of a second target category and inputting it into a second text encoder to generate a second similarity matrix; filtering image blocks of a second target category based on the second similarity matrix and a second discriminative category similarity and inputting them into a second image encoder; inputting the counting text into a third text encoder to generate a target similarity matrix; and obtaining the number of people in the image based on the similarity between the target similarity matrix and the counting text. This invention eliminates the need for any manual labeling, significantly reducing annotation costs.
Owner:HUAZHONG UNIV OF SCI & TECH

Cross-modal crowd counting method based on CNN and transformer

This invention discloses a cross-modal crowd counting method based on CNN and transformer. The method includes the following steps: inputting RGB images and thermal images into the branches of a dual-branch CNN network to learn modality-specific features of the dual-modal images; connecting the dual-branch CNN network with a novel cross-modal transformer to learn global features of different modal images, fusing modality-specific features and global modality features; connecting the fused feature maps from different layers of the network via a cross-layer connection structure, and enhancing the channel information of the fused feature maps through a branch attention module; extracting complementary information between different modalities using a cross-modal attention module to enhance cross-modal feature representation; feeding the feature maps extracted by the cross-modal attention module into a tail network to generate a density map; and summing the density maps pixel by pixel to obtain the crowd counting result. This invention can effectively complete cross-modal crowd counting tasks in crowded scenarios with arbitrary crowd distribution.
Owner:YANSHAN UNIV

A neural network-based crowd counting method and device

The application discloses a kind of crowd counting method and device based on neural network, method includes obtaining multiple sample images;Sample image is input into the original neural network model to be trained, and the first loss function of original neural network model is obtained;Sample image is cropped as the first subgraph and the second subgraph with overlapping area, and a subgraph and the second subgraph are respectively input into original neural network model, and the second loss function of original neural network model is obtained;According to the iterative training of first loss function and second loss function to original neural network model, to make the loss of original neural network model minimum, obtain target neural network model, to carry out crowd counting based on target neural network model.The application sets up auxiliary task, sample image is cropped as two subgraphs with overlapping area, and sample image and subgraph are input into network model and trained, to excavate the implicit relationship related to training set and counting task, improve the accuracy of network model prediction.
Owner:709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD