Congestion detection method and device based on multi-modal image
Through the crowding detection method of multimodal image fusion training, the accuracy and misjudgment problems of detection in high-density crowd scenes are solved, and efficient and accurate crowding detection is achieved in complex scenarios.
Patent Information
- Application Number
- CN202510534093.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing crowding detection methods are difficult to achieve accurate sub-region detection in high-density crowds or dense objects, and misjudgment is prone to changes in the sparseness of crowding.
The crowding detection method based on multimodal images is adopted, and the multimodal dependency relationship of graphic feature information is deeply explored, and the image and text feature vectors are fusion trained using the CLIP multimodal model to generate multimodal feature representations, and basic judgment parameters are set for detection.
It realizes accurate identification of crowded states of varying degrees in complex and variable crowding scenarios, improves detection accuracy and robustness, reduces the impact of occlusion factors, and maintains efficient detection performance.
Smart Images

Figure CN120279491A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and in particular to a crowd detection method and device based on multi-modal images. Background Art
[0002] In recent years, as a typical computer vision task, object detection technology has received increasing attention. With the continuous increase in crowd density and pedestrian flow, the crowd detection technology for high-density crowds has played an important role in many fields such as public safety, video surveillance, behavior understanding, and crowd evacuation management. Due to the uncertain distribution and occlusion problems of crowds, crowd detection has a high demand for the accuracy of local detection of images, and requires the algorithm to have both strong real-time performance and accurate detection results.
[0003] Currently, traditional crowd detection methods include methods based on object detection, which achieve the statistics of the number of people by identifying detected bounding boxes. Although such methods have accurate positioning, their recognition accuracy is insufficient, and the error rate is very high in crowded situations; moreover, precise crowd density detection in scenes with high-density crowds or dense objects is still a challenging task. The objects detected in dense scenes generally have local aggregation, which makes it difficult for the network to predict the crowd density in specific regions by region. Although such methods can verify the overall number of people, there are still limitations such as low recognition accuracy and uneven sparsity in practical applications.
[0004] Therefore, how to invent a crowd detection method based on multi-modal images that can identify and detect regions quickly and accurately has become an urgent problem to be solved. Summary of the Invention
[0005] To this end, the present invention provides a crowd detection method and device based on multi-modal images. By deeply exploring the multi-modal dependence relationship of text and image feature information and accurately modeling it, the feature expressions of different modal information are successfully distinguished, thereby realizing more accurate and reliable crowd density detection.
[0006] To achieve the above object, the present invention provides the following technical solution: A crowd detection method based on multi-modal images, including: Obtaining initial video data by collecting videos recorded by cameras in real subway scenes; constructing an image library by successively performing slicing, sorting, and cropping processing on the initial video data; Completing text information initialization by textually describing images with a set crowd degree and inputting the text description into a text library; Extract the features of the images in the image library through a set convolutional neural network to obtain image feature vectors; extract the features of the text information in the text library through a set text encoder to obtain text feature vectors; Fuse and train the image feature vectors and the text feature vectors through a CLIP multimodal model to generate multimodal feature representations; Set basic judgment parameters; based on the basic judgment parameters, detect and process the multimodal feature representations through a multimodal crowd detection model, and output detection results.
[0007] As a preferred solution of a crowd detection method based on multimodal images, during the process of text description of the images with the set crowding degree, set the types of text description to four types: low, medium, high, and ultra-high; the four types correspond to images with different crowding degrees from low to high in sequence.
[0008] As a preferred solution of a crowd detection method based on multimodal images, during the process of extracting the features of the text information in the text library through the set text encoder, calculate the cosine similarity between the feature vectors F X and F Y to judge the similarity between the image embedding F X and the corresponding text embedding F Y ; The calculation formula of the cosine similarity is:
[0009] In the formula, F X and F Y are feature vectors; is the dot product between vectors; and are the norms of vectors F X and vector F Y respectively.
[0010] As a preferred solution of a crowd detection method based on multimodal images, during the process of extracting the features of the text information in the text library through the set text encoder, quantify the relationship between positive and negative samples through a contrast loss function; the expression of the contrast loss function is:
[0011] In the formula, Loss NCE is the contrast loss function; , is the image feature vector in the positive sample; is the text feature in the negative sample; sim is the cosine similarity; is a parameter; is an indicator function, which is 1 when k≠x1, otherwise 0.
[0012] As a preferred solution of a multi-modal image-based crowd detection method, during the process of detecting and processing the multi-modal feature representation through the multi-modal crowd detection model, the classification result is optimized by a cross-entropy loss function; the expression of the cross-entropy loss function is:
[0013] In the formula, Loss CE is the cross-entropy loss function; N is the number of samples; M is the number of sample categories; is the label of sample i, taking 1 if the category is equal to c, otherwise taking 0; The probability that the i-th sample is predicted to belong to category c.
[0014] As a preferred solution of a multi-modal image-based crowd detection method, the basic judgment parameters include: crowding degree threshold, average distance between two targets, number of crowd blocks, number of people in a single crowd block area, and total number of people.
[0015] The present invention also provides a multi-modal image-based crowd detection device, based on the above multi-modal image-based crowd detection method, including: An image library construction module, configured to obtain initial video data by collecting videos recorded by cameras in a real subway scene; construct an image library by sequentially performing slicing, sorting, and cropping processing on the initial video data; A text information initialization module, configured to complete text information initialization by textually describing an image with a set crowding degree and inputting the text description into a text library; A feature vector extraction module, configured to extract image feature vectors from the images in the image library by setting a convolutional neural network; extract text feature vectors from the text information in the text library by setting a text encoder; A multi-modal feature representation generation module, configured to perform fusion training on the image feature vectors and the text feature vectors by a CLIP multi-modal model to generate a multi-modal feature representation; A multi-modal crowd detection model detection processing module, configured to set basic judgment parameters; based on the basic judgment parameters, detect and process the multi-modal feature representation by a multi-modal crowd detection model, and output a detection result.
[0016] As a preferred solution of a crowding detection device based on multi-modal images, in the text information initialization module, during the process of textually describing the images with the set crowding degree, the types of text descriptions are set to four types: low, medium, high, and ultra-high; the four types respectively correspond to images with different crowding degrees from low to high.
[0017] As a preferred solution of a crowding detection device based on multi-modal images, in the feature vector extraction module, during the process of extracting features from the text information in the text library through the set text encoder, the feature vectors F X and F Y of positive and negative sample pairs are calculated, and the similarity between the image embedding F X and the corresponding text embedding F Y is judged; The calculation formula for cosine similarity is:
[0018] In the formula, F X and F Y are feature vectors; is the dot product between vectors; and are respectively the norms of vectors F X and vector F Y respectively.
[0019] As a preferred solution of a crowding detection device based on multi-modal images, in the feature vector extraction module, during the process of extracting features from the text information in the text library through the set text encoder, the relationship between positive and negative samples is quantified through a contrastive loss function; the expression of the contrastive loss function is:
[0020] In the formula, Loss NCE is the contrastive loss function; , is the image feature vector in the positive sample; is the text feature in the negative sample; sim is the cosine similarity; is a parameter; is an indicator function, which is 1 when k≠x1, otherwise 0.
[0021] As a preferred solution of a crowding detection device based on multi-modal images, in the multi-modal crowding detection model detection and processing module, during the process of detecting and processing the multi-modal feature representation through the multi-modal crowding detection model, the classification result is optimized through a cross-entropy loss function; the expression of the cross-entropy loss function is:
[0022] In the formula, Loss CE is the cross-entropy loss function; N is the number of samples; M is the number of sample categories; is the label of sample i, taking 1 if the category is equal to c, otherwise taking 0; The probability that the i-th sample is predicted to belong to category c.
[0023] As a preferred solution of a multi-modal image-based crowd detection device, in the multi-modal crowd detection model detection and processing module, the basic judgment parameters include: crowding threshold, average distance between two targets, number of crowd blocks, number of people in a single crowd block area, and total number of people.
[0024] The present invention has the following advantages: The present invention obtains initial video data by collecting videos recorded by cameras in a real subway scene; constructs an image library by successively performing slicing, sorting, and cropping operations on the initial video data; completes text information initialization by performing text descriptions on images with a set degree of crowding and inputting the text descriptions into a text library; obtains image feature vectors by setting a convolutional neural network to extract features from the images in the image library; obtains text feature vectors by setting a text encoder to extract features from the text information in the text library; generates a multimodal feature representation by fusing and training the image feature vectors and the text feature vectors through a CLIP multimodal model; sets basic judgment parameters; and based on the basic judgment parameters, performs detection processing on the multimodal feature representation through a multimodal crowd detection model to output a detection result. By deeply exploring the multimodal dependence relationship of graphic and text feature information and accurately modeling it, the present invention successfully distinguishes the feature expressions of different modal information, thereby realizing more accurate and reliable crowding degree detection. On the one hand, the present invention ingeniously introduces a CLIP pre-training task. As an advanced cross-modal learning method, the core of CLIP lies in enabling the model to deeply understand and precisely align two different modal information, namely images and texts, through the learning of a large number of text-image pairs. The image encoder and the text encoder are endowed with rich semantic representation capabilities under the CLIP framework, which greatly promotes cross-modal understanding between language and vision. By utilizing this feature of CLIP, the present invention realizes the effective fusion of image and text information in the crowding detection task, further improving the accuracy and robustness of detection. On the other hand, aiming at the misjudgment problem that may be caused by different degrees of sparsity of the crowding degree, the present invention creatively proposes a crowding degree module. This module can accurately identify and distinguish different degrees of crowding states, thus effectively solving the misjudgment situation that may occur in traditional methods when the degree of crowding sparsity changes. This design enables the present invention to still maintain excellent detection performance in complex and changeable crowding scenes. The present invention is not affected by occlusion factors and can verify the effectiveness of crowding detection on a dense crowd dataset. Description of the Drawings
[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.
[0026] The structures, proportions, sizes, etc. shown in this specification are only used to match the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the implementation conditions of the present invention. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the efficacy that the present invention can produce and the purpose that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0027] Figure 1 It is a schematic flowchart of a crowd detection method based on multi-modal images provided in Embodiment 1 of the present invention; Figure 2 It is a specific implementation flowchart of a crowd detection method based on multi-modal images provided in Embodiment 1 of the present invention; Figure 3 It is a schematic diagram of four crowd situations of low, medium, high, and ultra-high in a possible embodiment provided in Embodiment 1 of the present invention; Figure 4 It is a schematic diagram of the architecture of a crowd detection device based on multi-modal images provided in Embodiment 2 of the present invention. Detailed implementation manners
[0028] The following specific embodiments illustrate the implementation manners of the present invention. Those familiar with this technology can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0029] Embodiment 1 Refer to Figure 1 and Figure 2 Embodiment 1 of the present invention provides a crowd detection method based on multi-modal images, including the following steps: S1. Obtain initial video data by collecting the video recorded by the camera in the real subway scene; construct an image library by successively performing slicing, sorting, and cropping processing on the initial video data; S2. Complete the initialization of text information by textually describing the images with a set crowd degree and inputting the text description into the text library; S3. Extract image feature vectors by setting a convolutional neural network for the images in the image library; extract text feature vectors by setting a text encoder for the text information in the text library; S4. Perform fusion training on the image feature vectors and the text feature vectors through a CLIP multi-modal model to generate multi-modal feature representations; S5. Set basic judgment parameters; based on the basic judgment parameters, detect and process the multi-modal feature representation through a multi-modal crowding detection model, and output a detection result.
[0030] In this embodiment, in step S1, initial video data is obtained by collecting the video recorded by the camera in the real subway scene; by sequentially performing slicing, sorting, and cropping on the initial video data, an image library is constructed. Specifically, initial video data is obtained by collecting the video recorded by the camera in the subway scene; first, the initial video data is split to ensure that the duration of each video segment is 30 seconds for subsequent processing. At the same time, the video picture quality is uniformly set to 1080p to ensure image clarity and detail presentation. Next, a strategy of collecting a single-frame image of the video every 5 seconds is adopted to construct an image library in this way. Each image in the image library has been standardized, and the size is uniformly set to 1024 × 1024 pixels to ensure the consistency of the image data.
[0031] In this embodiment, in step S2, the image with a set crowding degree is textually described, and the text description is input into the text library to complete the initialization of the text information. Specifically, the initialization of the text information focuses on corresponding the text descriptions corresponding to the images with different levels of crowding. First, the main types of text descriptions are set, which are divided into four types: low, medium, high, and ultra-high, and these four types correspond to the images with different levels of crowding from low to high in turn. Subsequently, the text description is input into the text library to complete the initialization of the text information, laying a foundation for the subsequent pre-training task.
[0032] In this embodiment, in step S3, an image feature vector is obtained by setting a convolutional neural network to extract features from the images in the image library; a text feature vector is obtained by setting a text encoder to extract features from the text information in the text library. Specifically, for image features, the convolutional neural network technology of the image encoder ResNet50 is used to extract features from the images in the video frames, generating feature vectors that can reflect the content and structure of the images. For text features, the text encoder Text Transformer is used. Text Transformer is responsible for converting the text information in the text library into a dense representation in a high-dimensional vector space to capture the semantic features of the text. The multi-modal extraction of images and texts provides richer and more comprehensive feature representations; since the feature extraction is performed in parallel, computing resources can be fully utilized to further improve the processing speed.
[0033] In this embodiment, during the process of feature extraction from the images in the image library, the ResNet50 network is responsible for the image feature extraction task. The ResNet50 network consists of 49 convolutional layers and 1 fully connected layer. The middle layers of the ResNet50 network are responsible for feature extraction, and the feature map of the image is output through the C4 layer of the middle layers of the ResNet50 network. The C4 layer consists of multiple residual blocks and skip connections. Each residual block consists of three convolutional layers, where the first 1x1 convolutional layer is used to reduce the number of channels, the second 3x3 convolutional layer is used to extract features, and the last 1x1 convolutional layer is used to restore the number of channels. The input of the image is X[n, h, w, c], where n represents the batch size; h represents the height of the image (height), which is 224 pixels; w represents the width of the image (width), which is 224 pixels; c represents the number of channels of the image (channel), and c = 3, indicating that this is the input of a color image.
[0034] The number of input images is N, N = 512; finally, N image features are output, denoted by F X = (F X1 , F X2 ,..., F XN ) set, and the output feature dimension is consistent with the input feature dimension.
[0035] In this embodiment, during the process of feature extraction from the text information in the text library, the crowd congestion level is classified into four intervals: [0, 10], [10, 20], [20, 50], and [50, 100], and the corresponding text information descriptions are four levels: low, medium, high, and extremely high. This is only applicable to the image data captured by a single camera in a fixed subway scenario. Then, the input of the text is tokenized as Y[n, l], where n is the batch size and l is the sequence length. Tokenization means converting the words in these four categories into a numerical form that the model can understand, that is, tokens. This is usually achieved by looking up a predefined vocabulary, and each word or sub-word in the vocabulary has a unique numerical ID. The dimension embedding size used when the text input is converted into a numerical vector is 512. This means that each text input is converted into a 512-dimensional vector.
[0036] The tokenized text is input into the text encoder Text Transformer to obtain the features of some text. Text information is simpler than image information. The text encoder Text Transformer consists of 12 layers of Transformer encoders. The Transformer encoder is composed of a self-attention module and a feed-forward neural network. The numhead value refers to the number of heads used in the multi-head self-attention mechanism in the Transformer model, and numhead = 8.
[0037] Also, since the image library and the text library are corresponding, the number of input texts is N, and finally N text features are output, represented by the set F Y = (F Y1 , F Y2 ,..., F YN ). The output feature dimension is the same as the input feature dimension. Finally, the text encoder will obtain 512 text features.
[0038] In this embodiment, in step S4, the CLIP multi-modal model is used to fuse and train the image feature vector and the text feature vector to generate a multi-modal feature representation; Specifically, the CLIP multi-modal model framework is introduced. CLIP uses a large number of image-text pairs for contrastive learning to learn the alignment relationship between images and texts. During the training process, the image and text features are bound together to form a unified multi-modal feature representation. This fusion method enables the model to utilize the information of both image and text modalities simultaneously, thereby improving the prediction accuracy and robustness. Through multi-modal fusion training, the model learns how to synthesize the features of different modalities to infer the congestion level of the subway, rather than relying solely on a single visual feature.
[0039] Among them, during the training process, the multi-modal model selects text-image pairs from the dataset and uses contrastive learning to distinguish positive and negative samples. Specifically, there are 64 text-image pairs in each batch. At this time, 64 images and 64 texts will be obtained simultaneously. First, a text-image pair is taken out from the 64 text-image pairs, and the paired text-image pair is a natural positive sample. And the other 63 images are all negative samples.
[0040] In this embodiment, by calculating the cosine similarity between the feature vectors F X and F Y of the positive and negative sample pairs, the similarity between the image embedding F X and the corresponding text embedding F Y is judged; The cosine similarity is calculated through the feature vector FX and F Y Perform a cross product between them to obtain a [64, 64] matrix. Then, the values on the diagonal are obtained from the pairwise feature inner products.
[0041] The calculation formula for cosine similarity is:
[0042] In the formula, F X and F Y are feature vectors; is the dot product between vectors; and are the norms of vectors F X and vector F Y respectively.
[0043] If the image embedding F X and the corresponding text embedding F Y are more similar, then the value of the cosine similarity is larger. Select the first row in the [64, 64] matrix, which represents the similarity degree between the 1st image and 64 texts. Among them, the 1st text is the positive sample, and set the label of this row to 1.
[0044] In this embodiment, the relationship between positive and negative samples is quantified through a contrast loss function; the expression of the contrast loss function is:
[0045] In the formula, Loss NCE is the contrast loss function; , is the image feature vector in the positive sample; is the text feature in the negative sample; sim is the cosine similarity; is a parameter; is an indicator function, which is 1 when k≠x1, otherwise 0.
[0046] In this embodiment, in step S5, set basic judgment parameters; based on the basic judgment parameters, detect and process the multi-modal feature representation through a multi-modal crowd detection model, and output a detection result.
[0047] Among them, the basic judgment parameters include: crowding degree threshold, average distance between two targets, number of crowd blocks, number of people in a single crowd block area, and total number of people.
[0048] Specifically, perform a preliminary judgment and classification on the target video image according to the crowding degree threshold. If the detected crowding intensity is not greater than the crowding intensity threshold, output that the target video image is empty; if the detected crowding intensity is greater than the crowding intensity threshold, calculate and output the number of people in the target video image and the corresponding crowding level.
[0049] When the detected crowding intensity is greater than the crowding intensity threshold, first determine the crowding degree of each crowd block; adaptively divide the target video image into blocks. The number of blocks is the number of equal divisions of the image, b = 8; calculate the average distance d between every two targets.
[0050] In the formula, are the coordinate positions of each person target in each block. By calculating the Euclidean distance between every two targets, it is determined whether the crowd is crowded.
[0051] According to the crowd target distance threshold (set to 0.5) and the block threshold (set to half of the number of blocks 8, which is 4), the overall crowding degree of the crowd, that is, the crowd distribution, can be determined. The distribution is also divided into 4 categories: multi-region dense distribution, single-region dense distribution, multi-region scattered distribution, and single-region scattered distribution. When the distance d > 0.5 in more than 4 blocks, it means the crowd belongs to the multi-region dense distribution. When the distance d > 0.5 in less than 4 blocks, it belongs to the single-region dense distribution. When the distance d ≤ 0.5 in more than 4 blocks, it means the crowd belongs to the multi-region scattered distribution. When the distance d ≤ 0.5 in less than 4 blocks, it means the crowd belongs to the single-region scattered distribution.
[0052] Sum up the number of people in the 8 blocks to get the total number of people S. S is divided into four intervals: [0, 10], [10, 20], [20, 50], [50, 100], corresponding to four levels: low, medium, high, and extremely high.
[0053]
[0054] In the formula, b is the number of blocks, which is 8; num is the number of people in a single block area.
[0055] According to the calculation results, the multi-modal crowding detection model outputs the number of people in the target video image and the corresponding crowding level.
[0056] In a possible embodiment, a specific experimental verification example is provided as follows: In this embodiment, in the subway platform and concourse, a large number of crowd flow videos and images are captured using high-definition cameras. These data cover different time periods (such as morning and evening rush hours, off-peak hours, etc.) and different types of crowding degrees to ensure the diversity and representativeness of the data set.
[0057] The collected data is labeled, and data augmentation is performed on the data images by means of rotation, translation, scaling, brightness or contrast adjustment, etc. Negative sample images related to the task are collected and appropriately labeled (marked as: no target category) to ensure a relatively balanced number of positive and negative samples in the training set. Finally, the labeled dataset is divided into a training set, a validation set, and a test set for subsequent model training and evaluation.
[0058] Through the method proposed in the present invention, the output of the congestion situation is as Figure 3 shown.
[0059] In this embodiment, the detection level of the model is evaluated by precision, recall, and the detection frame rate (speed FPS). After testing with the experimental dataset, the precision reaches 97.5%, the recall is 96%, and the detection frame rate is 27 FPS. The experimental results show that the present invention not only maintains high precision and recall but also has a relatively fast processing speed, meeting the real-time requirements.
[0060] In summary, the present invention obtains initial video data by collecting videos recorded by cameras in a real subway scenario; constructs an image library by successively performing slicing, sorting, and cropping operations on the initial video data; completes the initialization of text information by textually describing images of a set level of crowding and inputting the textual descriptions into a text library; obtains image feature vectors by setting a convolutional neural network to extract features from the images in the image library; obtains text feature vectors by setting a text encoder to extract features from the text information in the text library; generates a multimodal feature representation by fusing and training the image feature vectors and the text feature vectors through a CLIP multimodal model; sets basic judgment parameters; and based on the basic judgment parameters, performs detection processing on the multimodal feature representation through a multimodal crowd detection model to output a detection result. By deeply exploring the multimodal dependence relationship of graphic and text feature information and accurately modeling it, the present invention successfully differentiates the feature expressions of different modal information, thereby achieving more accurate and reliable crowding detection. On the one hand, the present invention ingeniously introduces the CLIP pre-training task. As an advanced cross-modal learning method, the core of CLIP lies in enabling the model to deeply understand and precisely align two different modal information, namely images and texts, through the learning of a large number of text-image pairs. The image encoder and the text encoder are endowed with rich semantic representation capabilities under the CLIP framework, which greatly promotes cross-modal understanding between language and vision. By leveraging this feature of CLIP, the present invention realizes the effective fusion of image and text information in the crowding detection task, further improving the accuracy and robustness of the detection. On the other hand, aiming at the misjudgment problem that may be caused by different degrees of crowding sparsity, the present invention creatively proposes a crowding degree module. This module can accurately identify and distinguish different levels of crowding states, thus effectively solving the misjudgment situation that may occur in traditional methods when the degree of crowding sparsity changes. This design enables the present invention to still maintain excellent detection performance in complex and variable crowding scenarios. The present invention is not affected by occlusion factors and can verify the effectiveness of crowding detection on a dense crowd dataset.
[0061] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0062] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0063] Embodiment 2 Referring to Figure 4 , Embodiment 2 of the present invention further provides a crowd detection device based on multimodal images, including: An image library construction module 001, configured to obtain initial video data by collecting videos recorded by cameras in a real subway scene; construct an image library by sequentially performing slicing, sorting, and cropping processing on the initial video data; A text information initialization module 002, configured to complete text information initialization by performing text descriptions on images with a set crowding degree and inputting the text descriptions into a text library; A feature vector extraction module 003, configured to extract image feature vectors by setting a convolutional neural network for images in the image library; extract text feature vectors by setting a text encoder for the text information in the text library; A multimodal feature representation generation module 004, configured to perform fusion training on the image feature vectors and the text feature vectors through a CLIP multimodal model to generate a multimodal feature representation; A multimodal crowd detection model detection processing module 005, configured to set basic judgment parameters; based on the basic judgment parameters, perform detection processing on the multimodal feature representation through a multimodal crowd detection model and output a detection result.
[0064] In this embodiment, in the text information initialization module 002, during the process of performing text descriptions on images with the set crowding degree, the types of text descriptions are set to four types: low, medium, high, and ultra-high; the four types correspond to images with different crowding degrees from low to high in sequence.
[0065] In this embodiment, in the feature vector extraction module 003, during the process of extracting the text information in the text library through the set text encoder, by calculating the cosine similarity between the feature vectors F X and F Y of positive and negative sample pairs, the image embedding F X and the corresponding text embedding F Yto determine the similarity; the calculation formula of cosine similarity is:
[0066] In the formula, F X and F Y are feature vectors; is the dot product between vectors; and are the norms of vectors F X and vector F Y respectively.
[0067] In this embodiment, in the feature vector extraction module 003, during the process of extracting features of the text information in the text library through the set text encoder, the relationship between positive and negative samples is quantified by a contrast loss function; the expression of the contrast loss function is:
[0068] In the formula, Loss NCE is the contrast loss function; , is the image feature vector in the positive sample; is the text feature in the negative sample; sim is the cosine similarity; is a parameter; is an indicator function, which is 1 when k≠x1, otherwise 0.
[0069] In this embodiment, in the multi-modal crowd detection model detection and processing module 005, during the process of detecting and processing the multi-modal feature representation through the multi-modal crowd detection model, the classification result is optimized by a cross-entropy loss function; the expression of the cross-entropy loss function is:
[0070] In the formula, Loss CE is the cross-entropy loss function; N is the number of samples; M is the number of sample categories; is the label of sample i, which is 1 if the category is equal to c, otherwise 0; is the probability that the i-th sample is predicted to belong to category c.
[0071] In this embodiment, in the multi-modal crowd detection model detection and processing module 005, the basic judgment parameters include: crowding degree threshold, average distance between two targets, number of crowd blocks, number of people in a single crowd block area, and total number of people.
[0072] It should be noted that for the information interaction, execution process, etc. among the modules of the above system, since they are based on the same concept as the method embodiments in Embodiment 1 of this application, the technical effects brought by them are the same as those of the method embodiments of this application. For the specific content, reference can be made to the description in the method embodiments shown above in this application, and details will not be repeated here.
[0073] Embodiment 3 Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which program code of a method for crowd detection based on multi-modal images is stored, and the program code includes instructions for executing a method for crowd detection based on multi-modal images according to Embodiment 1 or any possible implementation manner thereof.
[0074] The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrating one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0075] Embodiment 4 Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor; The processor and the memory communicate with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute a method for crowd detection based on multi-modal images according to Embodiment 1 or any possible implementation manner thereof by calling the program instructions.
[0076] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in the memory. The memory can be integrated in the processor or can exist independently outside the processor.
[0077] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.).
[0078] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing system. They can be concentrated on a single computing system or distributed on a network composed of multiple computing systems. Optionally, they can be implemented by program code executable by the computing system. Thus, they can be stored in the storage system and executed by the computing system. And in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules respectively, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0079] Although the present invention has been described in detail above with general descriptions and specific embodiments, on the basis of the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A crowd detection method based on multi-modal images, characterized in that, Including: Obtain initial video data by collecting videos recorded by cameras in a real subway scenario; Construct an image library by successively slicing, organizing, and cropping the initial video data; Complete the initialization of text information by textually describing images of a set crowding degree and inputting the text description into a text library; Extract image feature vectors by setting a convolutional neural network to extract features from the images in the image library; extract text feature vectors by setting a text encoder to extract features from the text information in the text library; Generate a multimodal feature representation by fusing and training the image feature vectors and the text feature vectors through a CLIP multimodal model; Set basic judgment parameters; based on the basic judgment parameters, perform detection processing on the multimodal feature representation through a multimodal crowding detection model and output a detection result.
2. The method for detecting crowding based on multimodal images according to claim 1, wherein, During the process of textually describing the images of the set crowding degree, set the types of text description to four types: low, medium, high, and extremely high; the four types correspond to images of different crowding degrees from low to high in sequence.
3. The method for detecting crowding based on multimodal images according to claim 2, wherein, In the process of extracting features from the text information in the text library through the set text encoder, the feature vectors F of the positive and negative sample pairs are calculated X and F Y The cosine similarity between them is used to judge the similarity between the image embedding F X and the corresponding text embedding F Y ; The calculation formula for cosine similarity is: ; Where, F X and F Y are eigenvectors; is the dot product between vectors; and are the norms of vectors F X and vector F Y respectively.
4. A method for detecting crowding based on multimodal images according to claim 3, characterized in that, During the process of extracting features from the text information in the text library by setting the text encoder, quantify the relationship between positive and negative samples through a contrast loss function; the expression of the contrast loss function is: ; where Loss NCE is the contrastive loss function; , is the image feature vector in the positive samples; is the text feature in the negative samples; sim is the cosine similarity; is a parameter; is an indicator function, which is 1 when k≠x1 and 0 otherwise.
5. The method for detecting crowding based on multi-modal images according to claim 4, wherein During the process of performing detection processing on the multimodal feature representation through the multimodal crowding detection model, optimize the classification result through a cross-entropy loss function; the expression of the cross-entropy loss function is: ; Where, Loss CE is the cross-entropy loss function; N is the number of samples; M is the number of sample categories; is the label of sample i, taking 1 if the category is equal to c, otherwise taking 0; is the probability that the i-th sample is predicted to belong to category c.
6. A method for detecting crowding based on multi-modal images according to claim 5, characterized in that, The basic judgment parameters include: crowding degree threshold, average distance between two targets, number of crowd blocks, number of people in a single crowd block area, and total number of people.
7. A crowdedness detection device based on multimodal images, which adopts a crowdedness detection method based on multimodal images according to any one of claims 1-6, characterized in that, Including: An image library construction module, which is used to obtain initial video data by collecting videos recorded by cameras in a real subway scenario; Construct an image library by successively slicing, organizing, and cropping the initial video data; A text information initialization module, which is used to complete the initialization of text information by textually describing images of a set crowding degree and inputting the text description into a text library; A feature vector extraction module, which is used to extract image feature vectors by setting a convolutional neural network to extract features from the images in the image library; extract text feature vectors by setting a text encoder to extract features from the text information in the text library; A multimodal feature representation generation module, which is used to generate a multimodal feature representation by fusing and training the image feature vectors and the text feature vectors through a CLIP multimodal model; A multimodal crowding detection model detection processing module, which is used to set basic judgment parameters; based on the basic judgment parameters, perform detection processing on the multimodal feature representation through a multimodal crowding detection model and output a detection result.
8. The crowdedness detection device based on multi-modal images according to claim 7, wherein, In the text information initialization module, during the process of textually describing the images of the set crowding degree, set the types of text description to four types: low, medium, high, and extremely high; the four types correspond to images of different crowding degrees from low to high in sequence.
9. The multi-modal-image-based overcrowding detection device according to claim 8, wherein, In the feature vector extraction module, during the process of extracting features of the text information in the text library through the set text encoder, the cosine similarity between the feature vectors F X and F Y is calculated to judge the similarity between the image embedding F X and the corresponding text embedding F Y ; The calculation formula of cosine similarity is as follows: ; Where, F X and F Y are eigenvectors; is the dot product between vectors; and are respectively the magnitudes of vectors F X and vector F Y 10. The multi-modal image-based crowd detection device according to claim 9, characterized in that, In the feature vector extraction module, during the process of extracting features of the text information in the text library through the set text encoder, the relationship between positive and negative samples is quantified through a contrastive loss function; the expression of the contrastive loss function is: ; where Loss NCE is the contrastive loss function; , is the image feature vector in the positive samples; is the text feature in the negative samples; sim is the cosine similarity; is the parameter; is the indicator function, which is 1 when k≠x1 and 0 otherwise.