Crowd re-identification method and device, terminal equipment and storage medium
By counting and height sorting of crowd image samples, and training of crowd analysis model and multimodal alignment model, the problem of low accuracy caused by manual labeling noise in the prior art is solved, and efficient crowd re-identification is achieved.
Patent Information
- Application Number
- CN202510177643.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
Existing population recognition technology relies on manual labeling boxes, resulting in noisy data and reducing the accuracy of the model.
By obtaining several crowd image samples for quantity annotation and height sorting annotation, the target annotation image is obtained, and input it into the crowd analysis model and the multimodal alignment model for iterative training until the loss function converges, and the trained crowd re-identification model is obtained.
This method can efficiently extract and analyze the crowd characteristics in the image to be identified, improve the accuracy of crowd re-identification, and adapt to different scenes and lighting conditions.
Smart Images

Figure CN120107890A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning model technology, and in particular to a crowd re-identification method, apparatus, terminal device and storage medium. Background Art
[0002] Crowd re-identification (G-ReID) aims to identify and track the identities of the same crowd members across multiple camera viewpoints, at different times, and in complex environments. Existing crowd re-identification technologies are mainly based on deep learning methods, which use neural networks to extract individual visual features from images and fuse these individual features into crowd features. Existing methods rely on manual annotation boxes to obtain the location of individuals, and require fine annotation of image data in advance to ensure that each pedestrian is accurately framed. In complex scenes, especially in dense crowds, severe occlusion, or large changes in lighting, the manually annotated rectangular boxes have a lot of noise. The noisy data causes the model to learn incorrect crowd feature representations, thereby reducing the accuracy of crowd re-identification. Summary of the invention
[0003] The embodiments of the present invention provide a method, apparatus, terminal device and storage medium for crowd re-identification, which can effectively solve the problem that the prior art relies on manually labeled boxes to obtain the location of individuals, the manually labeled rectangular boxes have a lot of noise, and the noisy data causes the model to learn incorrect crowd feature representations, resulting in low accuracy of crowd re-identification.
[0004] An embodiment of the present invention provides a method for re-identifying a crowd, including:
[0005] Obtain an image to be recognized;
[0006] The image to be identified is input into a preset crowd re-identification model to perform crowd feature recognition and obtain crowd re-identification results;
[0007] Wherein, the crowd re-identification model includes: a crowd analysis model and a multimodal alignment model;
[0008] The training of the crowd re-identification model includes:
[0009] Obtain several crowd image samples;
[0010] Perform quantity labeling and height sorting labeling according to the crowd image samples to obtain a target labeled image;
[0011] Inputting the target annotated image into the crowd analysis model to be trained for iterative training until the preset first loss function converges, thereby obtaining a trained crowd analysis model;
[0012] Inputting the crowd image samples into the trained crowd analysis model to perform crowd recognition, and obtaining crowd quantity characteristics and crowd sorting characteristics;
[0013] Inputting the crowd quantity feature and the crowd sorting feature into the multimodal alignment model to be trained for iterative training until the preset second loss function converges, thereby obtaining a trained multimodal alignment model;
[0014] According to the trained crowd analysis model and the multimodal alignment model, a trained crowd re-identification model is obtained.
[0015] Furthermore, the crowd analysis model includes: a crowd counting sub-model and a crowd sorting sub-model; the first loss function includes: a counting model loss function and a sorting model loss function;
[0016] The target annotated image is input into the crowd analysis model to be trained for iterative training until the preset first loss function converges, thereby obtaining a trained crowd analysis model, including:
[0017] Inputting the target annotated image into the crowd counting sub-model to be trained to extract features, thereby obtaining image modality features and text modality features;
[0018] Calculating a counting model loss function according to the image modality feature, the text modality feature, and the target annotated image;
[0019] When the loss function value of the counting model loss function converges, a trained crowd counting sub-model is obtained;
[0020] Inputting the image modality features and the text modality features into the crowd ranking sub-model to be trained for ranking prediction to obtain a predicted crowd ranking;
[0021] Calculating a ranking model loss function according to the predicted population ranking and the target annotated image;
[0022] When the loss function value of the ranking model loss function converges, a trained crowd ranking sub-model is obtained;
[0023] According to the trained crowd counting sub-model and the trained crowd sorting sub-model, a trained crowd analysis model is obtained.
[0024] Furthermore, the counting model loss function includes: a counting classification loss function and a counting contrast loss function;
[0025] Calculating a counting model loss function according to the image modality feature, the text modality feature, and the target annotated image includes:
[0026] Predicting based on the image modality features and the text modality features to obtain a predicted number of people;
[0027] Calculating a counting classification loss function according to the predicted number of people and the number annotation corresponding to the target annotated image;
[0028] generating, based on the target annotated image, a first counterfactual prompt for indicating a number of text annotation errors;
[0029] Calculating a count contrast loss function according to the first counterfactual prompt and the quantity annotation corresponding to the target annotated image;
[0030] The counting model loss function is calculated based on the counting classification loss function and the counting contrast loss function.
[0031] Furthermore, the ranking model loss function includes: a ranking classification loss function and a ranking comparison loss function;
[0032] Calculating a ranking model loss function according to the predicted crowd ranking and the target annotated image includes:
[0033] Calculate the ranking classification loss function according to the predicted population ranking and the height ranking annotation corresponding to the target annotated image;
[0034] generating a second counterfactual prompt for representing random ordering based on the target annotated image;
[0035] Calculate a ranking contrast loss function according to the second counterfactual prompt and the height ranking annotation corresponding to the target annotated image;
[0036] The sorting model loss function is calculated based on the sorting classification loss function and the sorting comparison loss function.
[0037] Furthermore, the crowd image samples are input into the trained crowd analysis model to perform crowd recognition, and crowd quantity characteristics and crowd sorting characteristics are obtained, including:
[0038] Inputting the crowd image sample into the trained crowd counting sub-model to identify the number of people, and obtaining the number characteristics of the crowd;
[0039] The crowd image samples are input into the trained crowd sorting sub-model to sort the crowd and obtain crowd sorting features.
[0040] Furthermore, the multimodal alignment model includes: a text editor and an image editor;
[0041] Inputting the crowd quantity feature and the crowd sorting feature into the multimodal alignment model to be trained for iterative training until the preset second loss function converges, thereby obtaining a trained multimodal alignment model, including:
[0042] Inputting the crowd quantity feature and the crowd sorting feature into the image editor to be trained to generate a mask, thereby obtaining a pedestrian mask;
[0043] The crowd quantity feature and the crowd sorting feature are spliced and input into a text editor to obtain text features;
[0044] Mapping the pedestrian mask and the text feature into a preset space, and calculating the cosine similarity;
[0045] Calculating a second loss function according to the cosine similarity and the target annotated image;
[0046] When the loss function value of the second loss function converges, a trained multimodal alignment model is obtained;
[0047] When the loss function value of the second loss function has not converged, the current parameters of the image editor are updated according to the second loss function, so that the image editor to be trained is updated according to the current parameters.
[0048] Furthermore, it also includes:
[0049] Obtaining image resolution and image noise of the crowd image sample;
[0050] Comparing the image resolution with a preset resolution threshold, and comparing the image noise with a preset noise threshold;
[0051] The crowd image samples corresponding to the image resolution being greater than a preset resolution threshold and the image noise being greater than a preset noise threshold are taken as the final crowd image samples.
[0052] As an improvement of the above solution, another embodiment of the present invention provides a crowd re-identification device, including:
[0053] An image acquisition module, used for acquiring an image to be identified;
[0054] The image re-identification module is used to input the image to be identified into a preset crowd re-identification model to perform crowd feature recognition and obtain crowd re-identification results;
[0055] A model training module, used for training the crowd re-identification model;
[0056] Wherein, the crowd re-identification model includes: a crowd analysis model and a multimodal alignment model;
[0057] The model training module includes:
[0058] An image sample acquisition submodule is used to acquire a number of crowd image samples;
[0059] An image annotation submodule is used to annotate the number and height of the crowd image samples to obtain a target annotated image;
[0060] A first training submodule is used to input the target annotated image into the crowd analysis model to be trained for iterative training until a preset first loss function converges to obtain a trained crowd analysis model;
[0061] A feature recognition submodule is used to input the crowd image sample into the trained crowd analysis model to perform crowd recognition and obtain crowd quantity features and crowd sorting features;
[0062] A second training submodule is used to input the population quantity feature and the population ranking feature into the multimodal alignment model to be trained for iterative training until the preset second loss function converges to obtain a trained multimodal alignment model;
[0063] The third training submodule is used to obtain a trained crowd re-identification model based on the trained crowd analysis model and the multimodal alignment model.
[0064] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, a crowd re-identification method as described in the above embodiment is implemented.
[0065] Another embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a crowd re-identification method described in the above embodiment.
[0066] By implementing the present invention, at least the following beneficial effects are achieved:
[0067] The present invention provides a crowd re-identification method, device, terminal equipment and storage medium. The method can obtain an image to be identified; input the image to be identified into a preset crowd re-identification model to perform crowd feature identification, and obtain a crowd re-identification result; wherein the crowd re-identification model includes: a crowd analysis model and a multimodal alignment model; training the crowd re-identification model includes: obtaining a number of crowd image samples; performing quantity annotation and height sorting annotation according to the crowd image samples to obtain a target annotated image; inputting the target annotated image into a crowd analysis model to be trained for iterative training until a preset first loss function converges to obtain a trained crowd analysis model; inputting the crowd image samples into the trained crowd analysis model for crowd identification to obtain crowd quantity features and crowd sorting features; inputting the crowd quantity features and the crowd sorting features into a multimodal alignment model to be trained for iterative training until a preset second loss function converges to obtain a trained multimodal alignment model; and obtaining a trained crowd re-identification model according to the trained crowd analysis model and the multimodal alignment model. Through quantity labeling and height sorting labeling, a crowd re-identification model including a crowd analysis model and a multimodal alignment model is trained to obtain a crowd re-identification model, which can efficiently extract and analyze the crowd features in the image to be identified. It does not require manual labeling to obtain the individual position. It only needs to iteratively train the crowd analysis model and the multimodal alignment model on the target labeled image after quantity labeling and height sorting labeling until the first loss function and the second loss function converge to obtain the crowd re-identification model, ensuring that the model has high accuracy in recognizing the number and sorting features of the crowd, thereby improving the accuracy of crowd re-identification; the crowd re-identification model trained by the target labeled image can adapt to crowd images taken in different scenes, different lighting conditions, different angles and distances, and integrate the crowd quantity characteristics and crowd sorting characteristics of the crowd image to further improve the accuracy of crowd re-identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 is a flow chart of a method for crowd re-identification provided by an embodiment of the present invention;
[0069] Figure 2 It is a structural schematic diagram of a crowd re-identification device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0070] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0071] See also Figure 1 , is a flow chart of a crowd re-identification method provided by an embodiment of the present invention, including:
[0072] S1, obtaining an image to be identified;
[0073] Specifically, the image to be recognized can be captured in real time by a camera or read from a file.
[0074] S2, inputting the image to be identified into a preset crowd re-identification model to perform crowd feature recognition and obtain crowd re-identification results;
[0075] Specifically, the image to be identified is input into the preset crowd re-identification model, the crowd features (high-dimensional vectors) are extracted, and the crowd re-identification results are output. It can accurately identify, sort and describe each person in a complex image, with strong robustness and generalization capabilities.
[0076] Wherein, the crowd re-identification model includes: a crowd analysis model and a multimodal alignment model;
[0077] Specifically, the crowd analysis model is used to identify the number of people in an image and the order of the crowd; the multimodal alignment model is used to perform text-image alignment on the image based on the number of people and the order of the crowd.
[0078] The training of the crowd re-identification model includes:
[0079] Step 1: Obtain several crowd image samples;
[0080] Preferably, it also includes:
[0081] Obtaining image resolution and image noise of the crowd image sample;
[0082] Comparing the image resolution with a preset resolution threshold, and comparing the image noise with a preset noise threshold;
[0083] The crowd image samples corresponding to the image resolution being greater than a preset resolution threshold and the image noise being greater than a preset noise threshold are taken as the final crowd image samples.
[0084] Specifically, the sources of crowd image samples include image acquisition devices such as surveillance cameras, mobile phone cameras, and the Internet. In order to improve the image quality, it is necessary to screen the collected image data, remove crowd images containing obvious noise or poor quality, and obtain the final crowd image samples. Poor quality images are usually manifested as a combination of multiple factors. Low-resolution images often lack sufficient details, resulting in blurred targets and difficulty in extracting effective features. Blurred or motion-blurred images lose clarity due to shooting jitter or rapid movement of the target, which also affects the recognition effect. In addition, overexposed or underexposed images can cause excessively high or low brightness, obscuring the details of the target and making it difficult to identify. Occlusion problems are also common, especially in crowded environments. If the target is partially or mostly occluded, complete feature information may not be obtained. Background interference is another important factor affecting image quality. Complex or dynamic backgrounds may be confused with the target, resulting in interference when extracting features. Low-contrast images also affect the distinction between the target and the background, making recognition difficult. Extreme shooting angles, such as looking down, looking up, or sideways, can also make the target's features unclear, thereby increasing the difficulty of recognition. Finally, image noise, especially in low light conditions, may cause unnecessary noise in the image, further reducing the image quality. Therefore, the above factors work together to cause poor image quality during crowd re-identification, affecting the recognition performance of the system. It is necessary to compare the image resolution with the preset resolution threshold, and the image noise with the preset noise threshold; and then take the crowd image sample corresponding to the image resolution greater than the preset resolution threshold and the image noise greater than the preset noise threshold as the final crowd image sample.
[0085] Step 2: perform quantity labeling and height sorting labeling according to the crowd image samples to obtain a target labeled image;
[0086] Specifically, the exact number of pedestrians in each picture is marked, and the order of pedestrians in each picture is marked from high to low according to their height, so as to obtain a target labeled image.
[0087] Step 3: input the target annotated image into the crowd analysis model to be trained for iterative training until the preset first loss function converges to obtain a trained crowd analysis model;
[0088] Specifically, the crowd analysis model includes: a crowd counting sub-model and a crowd sorting sub-model; the first loss function includes: a counting model loss function and a sorting model loss function;
[0089] The target annotated image is input into the crowd analysis model to be trained for iterative training until the preset first loss function converges, thereby obtaining a trained crowd analysis model, including:
[0090] Inputting the target annotated image into the crowd counting sub-model to be trained to extract features, thereby obtaining image modality features and text modality features;
[0091] Calculating a counting model loss function according to the image modality feature, the text modality feature, and the target annotated image;
[0092] When the loss function value of the counting model loss function converges, a trained crowd counting sub-model is obtained;
[0093] Inputting the image modality features and the text modality features into the crowd ranking sub-model to be trained for ranking prediction to obtain a predicted crowd ranking;
[0094] Calculating a ranking model loss function according to the predicted population ranking and the target annotated image;
[0095] When the loss function value of the ranking model loss function converges, a trained crowd ranking sub-model is obtained;
[0096] According to the trained crowd counting sub-model and the trained crowd sorting sub-model, a trained crowd analysis model is obtained.
[0097] Specifically, the crowd counting submodel is used to count the number of people in the crowd; the crowd ranking submodel is used to rank each person in the crowd. The counting model loss function and the ranking model loss function correspond to the loss functions of the crowd counting submodel and the crowd ranking submodel respectively. The counting model loss function measures the difference between the number of people predicted by the crowd counting submodel and the actual number of people; the ranking model loss function measures the difference between the predicted ranking of the crowd ranking submodel and the actual ranking.
[0098] In a preferred embodiment of the present invention, the image modality feature is a high-dimensional vector, which represents the overall visual features of the image, including the content, structure, local details of the image, and the semantic information expressed by the image. The text modality feature is also a high-dimensional vector, which captures the deep semantic information such as objects, actions, scenes, etc. described in the text. First, the target annotated image is input into the crowd counting sub-model to be trained to extract features, and the image modality features and the text modality features are obtained. Then, the counting model loss function is calculated according to the image modality features, the text modality features and the target annotated image; when the loss function value of the counting model loss function converges, the trained crowd counting sub-model is obtained, and when the loss function value of the counting model loss function does not converge, the parameters of the crowd counting sub-model are updated according to the counting model loss function to continue training until the loss function value of the counting model loss function converges. Then, the image modality features and the text modality features are input into the crowd sorting submodel to be trained for sorting prediction to obtain the predicted crowd sorting, and then the sorting model loss function is calculated according to the predicted crowd sorting and the target annotated image; when the loss function value of the sorting model loss function converges, the trained crowd sorting submodel is obtained, and when the loss function value of the sorting model loss function does not converge, the parameters of the crowd sorting submodel are updated according to the sorting model loss function until the loss function value of the sorting model loss function converges. Finally, according to the trained crowd counting submodel and the trained crowd sorting submodel, the trained crowd analysis model is obtained. By continuously optimizing the parameters of the two submodels to minimize their respective loss function values until they converge, the trained crowd analysis model can accurately estimate the number of people in the image and effectively sort the crowd.
[0099] Specifically, the counting model loss function includes: a counting classification loss function and a counting contrast loss function;
[0100] Calculating a counting model loss function according to the image modality feature, the text modality feature, and the target annotated image includes:
[0101] Predicting based on the image modality features and the text modality features to obtain a predicted number of people;
[0102] Calculating a counting classification loss function according to the predicted number of people and the number annotation corresponding to the target annotated image;
[0103] generating, based on the target annotated image, a first counterfactual prompt for indicating a number of text annotation errors;
[0104] Calculating a count contrast loss function according to the first counterfactual prompt and the quantity annotation corresponding to the target annotated image;
[0105] The counting model loss function is calculated based on the counting classification loss function and the counting contrast loss function.
[0106] Specifically, the count classification loss function is used to measure the difference between the number of people predicted by the model and the actual number of people in the target annotated image, and is used to quantify the deviation between the predicted value and the actual value. The count contrast loss function is a loss function based on counterfactual prompts, which is used to deal with possible errors in text annotations. Counterfactual prompts are hypothetical information used to indicate how different the results would be if certain conditions were changed. In this embodiment, the first counterfactual prompt is used to represent the hypothetical number of text annotation errors.
[0107] In a preferred embodiment of the present invention, firstly, prediction is performed according to the image modality feature and the text modality feature to obtain the predicted number of people; then, according to the predicted number of people and the number annotation corresponding to the target annotation image, the counting classification loss function is calculated; then, according to the target annotation image, a first counterfactual prompt for indicating the number of text annotation errors is generated, that is, according to the target annotation image, the possible errors in the text annotation are analyzed, and the first counterfactual prompt for indicating the number of text annotation errors is generated, which may be a hypothetical quantity value, indicating how many people should be if the text annotation is correct. Then, according to the first counterfactual prompt and the number annotation corresponding to the target annotation image, the counting contrast loss function is calculated; finally, according to the counting classification loss function and the counting contrast loss function, the counting model loss function is calculated. In this embodiment, the counting classification loss function and the counting contrast loss function can be combined, or the final counting model loss function can be obtained by weighted summation. By introducing the counting contrast loss function, the counting model can not only measure the difference between the predicted value and the actual value, but also handle the possible errors in the text annotation, thereby improving the robustness and accuracy of the model.
[0108] Schematically, the ranking model loss function includes: a ranking classification loss function and a ranking comparison loss function;
[0109] Calculating a ranking model loss function according to the predicted crowd ranking and the target annotated image includes:
[0110] Calculate the ranking classification loss function according to the predicted population ranking and the height ranking annotation corresponding to the target annotated image;
[0111] generating a second counterfactual prompt for representing random ordering based on the target annotated image;
[0112] Calculate a ranking contrast loss function according to the second counterfactual prompt and the height ranking annotation corresponding to the target annotated image;
[0113] The sorting model loss function is calculated based on the sorting classification loss function and the sorting comparison loss function.
[0114] Specifically, the ranking classification loss function is used to measure the difference between the crowd ranking predicted by the model and the actual height ranking in the target annotated image, and is used to quantify the inconsistency between the predicted ranking and the actual ranking. The ranking contrast loss function is a loss function based on counterfactual prompts, which is used to deal with hypothetical ranking situations, that is, how different the results will be if the crowd is sorted in a random way. In this embodiment, the second counterfactual prompt is used to represent the hypothetical situation of this random ranking.
[0115] In a preferred embodiment of the present invention, a sorting classification loss function is first calculated based on the predicted crowd sorting and the height sorting annotation corresponding to the target annotated image; then, a second counterfactual prompt for representing random sorting is generated based on the target annotated image, which is used to simulate the situation where the crowd is not sorted by height but randomly sorted; then, a sorting comparison loss function is calculated based on the second counterfactual prompt and the height sorting annotation corresponding to the target annotated image; finally, a sorting model loss function is calculated based on the sorting classification loss function and the sorting comparison loss function by combining or weighted summing.
[0116] Then, the crowd sorting model is trained. First, the annotated and sorted images and texts are input into the pre-trained large model CLIP. The features of the text and image are extracted respectively through its text encoder and image encoder, and the corresponding predicted sorting is generated in the shared embedding space. Then, the sorting classification loss function is calculated based on the input annotation sorting and the predicted sorting generated by the model, encouraging the model to embed the image close to its correct text annotation. At the same time, an incorrect random sorting is generated for each training sample as a counterfactual prompt. By calculating the sorting contrast loss function, the image embedding is kept away from the text embedding containing the incorrect sorting, thereby constraining the distinguishability between the correct sorting and the counterfactual sorting. The training process continuously optimizes the model parameters through backpropagation, while minimizing the sorting classification loss and the sorting contrast loss until both loss functions converge, and finally completes the training of the model, enabling it to accurately sort pedestrians in the image.
[0117] In another preferred embodiment of the present invention, the similarity between the image and the correct text annotation (including the accurate number) is calculated based on the feature vectors of the image modality features and the text modality features to obtain the predicted count of the model. Then, the count classification loss function is calculated, which measures the accuracy of the model according to the difference between the predicted count and the annotation count. In addition, in order to enhance the robustness of the model, the system generates a counterfactual prompt for each training sample, that is, the number of errors contained in the text annotation, and inputs it into the model. The counterfactual prompt is automatically created by replacing the real object count in the original prompt with the error count. By calculating the count contrast loss function, the model is encouraged to embed the image close to the text annotation of the correct number, while pushing it away from the text annotation of the wrong number. During the training process, the model will continuously optimize the two loss functions, adjust the distance between the image and text features, so that the image features are not only close to the correct text annotation, but also pull the embedding distance from the counterfactual prompt, and finally achieve sensitivity to the number and improve the counting accuracy of the model. The training continues until the two loss functions converge, marking the completion of the training of the crowd counting model.
[0118] Step 4: Input the crowd image sample into the trained crowd analysis model to perform crowd recognition, and obtain crowd quantity characteristics and crowd sorting characteristics;
[0119] Specifically, the crowd image sample is input into the trained crowd counting sub-model to identify the number of people, and obtain the crowd number feature;
[0120] The crowd image samples are input into the trained crowd sorting sub-model to sort the crowd and obtain crowd sorting features.
[0121] Step 5: inputting the crowd quantity feature and the crowd sorting feature into the multimodal alignment model to be trained for iterative training until the preset second loss function converges, thereby obtaining a trained multimodal alignment model;
[0122] Preferably, the multimodal alignment model includes: a text editor and an image editor;
[0123] Inputting the crowd quantity feature and the crowd sorting feature into the multimodal alignment model to be trained for iterative training until the preset second loss function converges, thereby obtaining a trained multimodal alignment model, including:
[0124] Inputting the crowd quantity feature and the crowd sorting feature into the image editor to be trained to generate a mask, thereby obtaining a pedestrian mask;
[0125] The crowd quantity feature and the crowd sorting feature are spliced and input into a text editor to obtain text features;
[0126] Mapping the pedestrian mask and the text feature into a preset space, and calculating the cosine similarity;
[0127] Calculating a second loss function according to the cosine similarity and the target annotated image;
[0128] When the loss function value of the second loss function converges, a trained multimodal alignment model is obtained;
[0129] When the loss function value of the second loss function has not converged, the current parameters of the image editor are updated according to the second loss function, so that the image editor to be trained is updated according to the current parameters.
[0130] Specifically, the text editor is used to process and generate text features related to the crowd, which may include descriptions of the number of people, attributes of crowd sorting, etc. The image editor is used to process and generate visual features related to the crowd image, especially masks that can highlight pedestrians. In the image editor, the input crowd number features and crowd sorting features are used to generate a pedestrian mask. The pedestrian mask may be a binary image, in which the white area represents the pedestrian and the black area represents the background. For example, the semantic segmentation model SAM is used for automatic recognition to generate a mask corresponding to each person. The pedestrian mask is extracted by bitwise operation to extract the area in the image where the corresponding mask is 1, that is, the part where the pedestrian is located, and the background is removed; then, the extracted pedestrian area is subjected to necessary preprocessing, such as cropping and resizing, so as to extract its features. In this way, the mask helps to effectively focus on the pedestrian area and remove irrelevant background, thereby extracting more accurate and useful features. Then, based on the crowd counting and crowd sorting, the crowd number features and the crowd sorting features are spliced and input into the text editor to obtain text features. Specifically, in order to handle the difference in the number of pedestrians in different images, this embodiment designs an upper limit on the number of pedestrians in each group. For images with insufficient number of pedestrians, zeros are added to the end of the feature vector to make it reach the maximum length and ensure the consistency of the input feature dimension. At the same time, this embodiment uses the CoOp method to generate a corresponding implicit text description for each pedestrian mask, and obtains text features through a text encoder.
[0131] Then, the pedestrian mask and the text feature are mapped to a preset space, and the cosine similarity is calculated; the second loss function is calculated according to the cosine similarity and the target annotated image; when the loss function value of the second loss function converges, a trained multimodal alignment model is obtained; when the loss function value of the second loss function does not converge, the current parameters of the image editor are updated according to the second loss function, so that the image editor to be trained is updated according to the current parameters. In this process, the parameters of the text encoder of the multimodal alignment model (such as CLIP, Contrastive Language-Image Pre-training) are frozen, and only the parameters of the image encoder are updated. The crowd counting submodel and the crowd sorting submodel are used to obtain the crowd quantity text prompt and the sorting prompt respectively, and the two prompts are spliced and input into the text encoder of CLIP to obtain the text feature (high-dimensional vector), which is combined with the pedestrian feature (high-dimensional vector) obtained by the mask and mapped to a common embedding space. The similarity between the text and the image is measured by calculating the cosine similarity between the features. Then, a contrastive learning method is used to calculate the second loss function based on the cosine similarity and the target annotated image. The second loss function is used to optimize the similarity between the image and the text in the embedding space to ensure that the related images and texts are close in space and the unrelated images and texts are far away. Specifically, the similarity of positive samples (matching image and text pairs) is maximized and the similarity of negative samples (unmatched image and text pairs) is minimized, thereby achieving effective text-image alignment. Finally, when the loss function value of the second loss function converges, a trained multimodal alignment model is obtained; when the loss function value of the second loss function does not converge, the current parameters of the image editor are updated according to the second loss function, so that the image editor to be trained is updated according to the current parameters to obtain a trained image encoder, thereby obtaining a trained multimodal alignment model.
[0132] Step 6: Based on the trained crowd analysis model and the multimodal alignment model, a trained crowd re-identification model is obtained.
[0133] In a preferred embodiment of the present invention, by combining the powerful cross-modal learning ability of the CLIP model, contrastive learning is used to optimize the alignment of images and texts in a common embedding space, which can efficiently achieve crowd counting, sorting and individual recognition. By automatically generating counterfactual prompts and loss functions, the model can accurately distinguish between correct counting and sorting and incorrect counterfactual situations, thereby improving the sensitivity and accuracy of crowd details. At the same time, the SAM model is used for human mask segmentation, which further improves the accuracy of individual feature extraction. Ultimately, the method can accurately identify, sort and describe each person in a complex image, with strong robustness and generalization capabilities.
[0134] By implementing this embodiment, an image to be identified is obtained; the image to be identified is input into a preset crowd re-identification model to perform crowd feature identification, and a crowd re-identification result is obtained; wherein the crowd re-identification model includes: a crowd analysis model and a multimodal alignment model; the training of the crowd re-identification model includes: obtaining a number of crowd image samples; performing quantity annotation and height ranking annotation according to the crowd image samples to obtain a target annotated image; inputting the target annotated image into the crowd analysis model to be trained for iterative training until a preset first loss function converges to obtain a trained crowd analysis model; inputting the crowd image samples into the trained crowd analysis model for crowd identification to obtain crowd quantity features and crowd ranking features; inputting the crowd quantity features and the crowd ranking features into the multimodal alignment model to be trained for iterative training until a preset second loss function converges to obtain a trained multimodal alignment model; and obtaining a trained crowd re-identification model based on the trained crowd analysis model and the multimodal alignment model. Through quantity labeling and height sorting labeling, a crowd re-identification model including a crowd analysis model and a multimodal alignment model is trained to obtain a crowd re-identification model, which can efficiently extract and analyze the crowd features in the image to be identified. It does not require manual labeling to obtain the individual position. It only needs to iteratively train the crowd analysis model and the multimodal alignment model on the target labeled image after quantity labeling and height sorting labeling until the first loss function and the second loss function converge to obtain the crowd re-identification model, ensuring that the model has high accuracy in recognizing the number and sorting features of the crowd, thereby improving the accuracy of crowd re-identification; the crowd re-identification model trained by the target labeled image can adapt to crowd images taken in different scenes, different lighting conditions, different angles and distances, and integrate the crowd quantity characteristics and crowd sorting characteristics of the crowd image to further improve the accuracy of crowd re-identification.
[0135] See also Figure 2 , is a schematic diagram of the structure of a crowd re-identification device provided by an embodiment of the present invention, comprising:
[0136] An image acquisition module, used for acquiring an image to be identified;
[0137] The image re-identification module is used to input the image to be identified into a preset crowd re-identification model to perform crowd feature recognition and obtain crowd re-identification results;
[0138] A model training module, used for training the crowd re-identification model;
[0139] Wherein, the crowd re-identification model includes: a crowd analysis model and a multimodal alignment model;
[0140] The model training module includes:
[0141] An image sample acquisition submodule is used to acquire a number of crowd image samples;
[0142] An image annotation submodule is used to annotate the number and height of the crowd image samples to obtain a target annotated image;
[0143] A first training submodule is used to input the target annotated image into the crowd analysis model to be trained for iterative training until a preset first loss function converges to obtain a trained crowd analysis model;
[0144] A feature recognition submodule is used to input the crowd image sample into the trained crowd analysis model to perform crowd recognition and obtain crowd quantity features and crowd sorting features;
[0145] A second training submodule is used to input the population quantity feature and the population ranking feature into the multimodal alignment model to be trained for iterative training until the preset second loss function converges to obtain a trained multimodal alignment model;
[0146] The third training submodule is used to obtain a trained crowd re-identification model based on the trained crowd analysis model and the multimodal alignment model.
[0147] The present invention provides a crowd re-identification device, which obtains an image to be identified according to an image acquisition module; inputs the image to be identified into a preset crowd re-identification model through the image re-identification module to perform crowd feature recognition to obtain a crowd re-identification result; a model training module is used to train the crowd re-identification model; wherein the crowd re-identification model includes: a crowd analysis model and a multimodal alignment model; obtains a plurality of crowd image samples through an image sample acquisition submodule; performs quantity annotation and height sorting annotation according to the crowd image samples in an image annotation submodule to obtain a target annotation image; inputs the target annotation image into a crowd re-identification model to be trained through a first training submodule; Iterative training is performed in the group analysis model until the preset first loss function converges to obtain a trained crowd analysis model; according to the feature recognition submodule, the crowd image samples are input into the trained crowd analysis model for crowd recognition to obtain crowd quantity characteristics and crowd sorting characteristics; then in the second training submodule, the crowd quantity characteristics and the crowd sorting characteristics are input into the multimodal alignment model to be trained for iterative training until the preset second loss function converges to obtain a trained multimodal alignment model; finally, in the third training submodule, a trained crowd re-identification model is obtained based on the trained crowd analysis model and the multimodal alignment model. Through quantity labeling and height sorting labeling, a crowd re-identification model including a crowd analysis model and a multimodal alignment model is trained to obtain a crowd re-identification model, which can efficiently extract and analyze the crowd features in the image to be identified. It does not require manual labeling to obtain the individual position. It only needs to iteratively train the crowd analysis model and the multimodal alignment model on the target labeled image after quantity labeling and height sorting labeling until the first loss function and the second loss function converge to obtain the crowd re-identification model, ensuring that the model has high accuracy in recognizing the number and sorting features of the crowd, thereby improving the accuracy of crowd re-identification; the crowd re-identification model trained by the target labeled image can adapt to crowd images taken in different scenes, different lighting conditions, different angles and distances, and integrate the crowd quantity characteristics and crowd sorting characteristics of the crowd image to further improve the accuracy of crowd re-identification.
[0148] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.
[0149] Those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0150] Another embodiment of the present invention further provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements a crowd re-identification method as described in the above embodiment when executing the computer program. The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.
[0151] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire terminal device.
[0152] The memory can be used to store the computer program, and the processor realizes various functions of the terminal device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device or other volatile solid-state storage device.
[0153] Another embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a crowd re-identification method described in the above embodiment.
[0154] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium.
[0155] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A crowd re-identification method, characterized in that: include: Obtain an image to be recognized; The image to be identified is input into a preset crowd re-identification model to perform crowd feature recognition and obtain crowd re-identification results; Wherein, the crowd re-identification model includes: a crowd analysis model and a multimodal alignment model; The training of the crowd re-identification model includes: Obtain several crowd image samples; Perform quantity labeling and height sorting labeling according to the crowd image samples to obtain a target labeled image; Inputting the target annotated image into the crowd analysis model to be trained for iterative training until the preset first loss function converges, thereby obtaining a trained crowd analysis model; Inputting the crowd image samples into the trained crowd analysis model to perform crowd recognition, and obtaining crowd quantity characteristics and crowd sorting characteristics; Inputting the crowd quantity feature and the crowd sorting feature into the multimodal alignment model to be trained for iterative training until the preset second loss function converges, thereby obtaining a trained multimodal alignment model; According to the trained crowd analysis model and the multimodal alignment model, a trained crowd re-identification model is obtained.
2. A method for crowd re-identification according to claim 1, characterized in that: The crowd analysis model includes: a crowd counting sub-model and a crowd sorting sub-model; the first loss function includes: a counting model loss function and a sorting model loss function; The target annotated image is input into the crowd analysis model to be trained for iterative training until the preset first loss function converges, thereby obtaining a trained crowd analysis model, including: Inputting the target annotated image into the crowd counting sub-model to be trained to extract features, thereby obtaining image modality features and text modality features; Calculating a counting model loss function according to the image modality feature, the text modality feature, and the target annotated image; When the loss function value of the counting model loss function converges, a trained crowd counting sub-model is obtained; Inputting the image modality features and the text modality features into the crowd ranking sub-model to be trained for ranking prediction to obtain a predicted crowd ranking; Calculating a ranking model loss function according to the predicted population ranking and the target annotated image; When the loss function value of the ranking model loss function converges, a trained crowd ranking sub-model is obtained; According to the trained crowd counting sub-model and the trained crowd sorting sub-model, a trained crowd analysis model is obtained.
3. A method for crowd re-identification according to claim 2, characterized in that: The counting model loss function includes: a counting classification loss function and a counting contrast loss function; Calculating a counting model loss function according to the image modality feature, the text modality feature, and the target annotated image includes: Predicting based on the image modality features and the text modality features to obtain a predicted number of people; Calculating a counting classification loss function according to the predicted number of people and the number annotation corresponding to the target annotated image; generating, based on the target annotated image, a first counterfactual prompt for indicating a number of text annotation errors; Calculating a count contrast loss function according to the first counterfactual prompt and the quantity annotation corresponding to the target annotated image; The counting model loss function is calculated based on the counting classification loss function and the counting contrast loss function.
4. A method for crowd re-identification according to claim 2, characterized in that: The ranking model loss function includes: a ranking classification loss function and a ranking comparison loss function; Calculating a ranking model loss function according to the predicted crowd ranking and the target annotated image includes: Calculate the ranking classification loss function according to the predicted population ranking and the height ranking annotation corresponding to the target annotated image; generating a second counterfactual prompt for representing random ordering based on the target annotated image; Calculate a ranking contrast loss function according to the second counterfactual prompt and the height ranking annotation corresponding to the target annotated image; The sorting model loss function is calculated based on the sorting classification loss function and the sorting comparison loss function.
5. A method for crowd re-identification according to claim 2, characterized in that: The crowd image samples are input into the trained crowd analysis model to perform crowd identification, and crowd quantity characteristics and crowd sorting characteristics are obtained, including: Inputting the crowd image sample into the trained crowd counting sub-model to identify the number of people, and obtaining the crowd number characteristics; The crowd image samples are input into the trained crowd sorting sub-model to sort the crowd and obtain crowd sorting features.
6. A method for crowd re-identification according to claim 1, characterized in that: The multimodal alignment model includes: a text editor and an image editor; Inputting the crowd quantity feature and the crowd sorting feature into the multimodal alignment model to be trained for iterative training until the preset second loss function converges, thereby obtaining a trained multimodal alignment model, including: Inputting the crowd quantity feature and the crowd sorting feature into the image editor to be trained to generate a mask, thereby obtaining a pedestrian mask; The crowd quantity feature and the crowd sorting feature are spliced and input into a text editor to obtain text features; Mapping the pedestrian mask and the text feature into a preset space, and calculating the cosine similarity; Calculating a second loss function according to the cosine similarity and the target annotated image; When the loss function value of the second loss function converges, a trained multimodal alignment model is obtained; When the loss function value of the second loss function has not converged, the current parameters of the image editor are updated according to the second loss function, so that the image editor to be trained is updated according to the current parameters.
7. A method for crowd re-identification according to claim 1, characterized in that: Also includes: Obtaining image resolution and image noise of the crowd image sample; Comparing the image resolution with a preset resolution threshold, and comparing the image noise with a preset noise threshold; The crowd image samples corresponding to the image resolution being greater than a preset resolution threshold and the image noise being greater than a preset noise threshold are taken as the final crowd image samples.
8. A crowd re-identification device, characterized in that: include: An image acquisition module, used for acquiring an image to be identified; The image re-identification module is used to input the image to be identified into a preset crowd re-identification model to perform crowd feature recognition and obtain crowd re-identification results; A model training module, used for training the crowd re-identification model; Wherein, the crowd re-identification model includes: a crowd analysis model and a multimodal alignment model; The model training module includes: An image sample acquisition submodule is used to acquire a number of crowd image samples; An image annotation submodule is used to annotate the number and height of the crowd image samples to obtain a target annotated image; A first training submodule is used to input the target annotated image into the crowd analysis model to be trained for iterative training until a preset first loss function converges to obtain a trained crowd analysis model; A feature recognition submodule is used to input the crowd image sample into the trained crowd analysis model to perform crowd recognition and obtain crowd quantity features and crowd sorting features; A second training submodule is used to input the population quantity feature and the population ranking feature into the multimodal alignment model to be trained for iterative training until the preset second loss function converges to obtain a trained multimodal alignment model; The third training submodule is used to obtain a trained crowd re-identification model based on the trained crowd analysis model and the multimodal alignment model.
9. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a method for crowd re-identification as claimed in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute a crowd re-identification method as claimed in any one of claims 1 to 7.
Citation Information
Cited By
Community personnel multi-mode identification system and method based on image perception large model
CN120726674A