A person re-identification method, device and electronic device
By building a character feature library in the monitoring system and using vector search technology, combining details and overall feature extraction methods, the problems of low efficiency and misidentification of specific characters in the existing technology are solved, and efficient and accurate character feature retrieval is achieved.
Patent Information
- Application Number
- CN202111362132.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-11-17
AI Technical Summary
The prior art is inefficient in finding specific characters in video surveillance, and there are problems of misidentification and misidentification, especially the method based on convolutional neural networks, resulting in insufficient feature extraction.
The vector search technology is used to construct a monitoring feature library in advance, convert the character images taken by the camera into character feature vectors, and store them in a classified manner. The feature accuracy is improved through the details and overall feature extraction methods, and the correlation relationship between the details feature vector and the overall feature vector is used, and the Milvus massive feature storage service is used for rapid search.
It improves the efficiency and accuracy of finding target characters in the monitoring system, reduces misidentification and misidentification, and realizes real-time and efficient character feature retrieval.
Smart Images

Figure CN114092881B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of visual image processing. Specifically, it relates to a person re-identification method, device, and electronic device. Background Art
[0002] With the continuous development of society, currently in the market, video surveillance technology and systems have been widely applied. For example, a certain number of cameras have been installed in office areas, squares, banks, streets, etc., realizing the functions of real-time collection, viewing, and storage of surveillance videos in each area. With the popularization of video surveillance, it has brought certain conveniences to all walks of life. For example: First, in the field of security, it has, to a certain extent, prevented the occurrence of bad behaviors (theft) and reduced the technical barriers for investigators to break through bad events (theft); Second, for places where personnel need to patrol, the responsible personnel can also confirm whether the patrol personnel have carried out regular inspections of specific equipment on time through the video surveillance recordings. However, for both of the above situations, it is necessary for the responsible personnel to check the video surveillance recordings one by one, which is a heavy task, not only time-consuming and laborious, but also inefficient.
[0003] With the continuous development of technology, there have also emerged some products on the market that automatically search for suspicious persons or the trajectories of specific patrol personnel from video recordings of different cameras. For the monitoring system, it is necessary to analyze and compare a large amount of images, and the efficiency of finding specific persons is low; moreover, most of them use a single convolutional neural network (CNN) technology to extract person features and search for specific persons through brute-force matching. The disadvantages of this technical route are: due to the inductive bias of CNN, the extracted person features are relatively one-sided, which easily causes a large number of misidentifications and missed identifications, and at the same time, the time wasted by brute-force matching. The above defects are not conducive to the large-scale promotion and use of existing products. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a person re-identification method, device, electronic device, and medium. Based on vector retrieval technology, it improves the efficiency of finding a specific target person from camera recordings, extracts more accurate person features, avoids the low distinguishability of the features extracted in the prior art, and enhances the accuracy of specific person retrieval.
[0005] A person re-identification method provided by an embodiment of the present application is applied to a monitoring system, and the monitoring system includes at least one camera and a pre-constructed monitoring feature library; the monitoring feature library includes at least one person feature library; each person feature library corresponds to a camera, and the person feature library includes the attribute information of the camera and the monitoring features of each person image captured by the camera, wherein the monitoring feature of each person image corresponds to the shooting time information of the person image; the method includes:
[0006] Obtain an image of a target person, and extract a person feature vector of the target person from the image of the target person as a feature to be retrieved;
[0007] For the feature to be retrieved, determine whether there is a monitoring feature in each person feature library in the monitoring feature library whose similarity with the feature to be retrieved satisfies a first preset condition;
[0008] If there is, determine an identification result according to the monitoring feature whose similarity satisfies the first preset condition, and the identification result includes the shooting time information corresponding to the monitoring feature that satisfies the first preset condition, and the camera attribute information in the person feature library where the monitoring feature that satisfies the first preset condition exists.
[0009] In some embodiments, in the person re-identification method, when extracting a person feature vector of a target person from the image of the target person as a feature to be retrieved, the extraction method includes:
[0010] For the image of the target person, extract a detail feature vector of the image of the target person, and the detail feature vector represents the detail information of multiple parts of the target person;
[0011] For the detail feature vector, extract an overall feature vector of the image of the target person, and the overall feature vector represents the overall feature information of the target person, and the overall feature information includes the detail information of multiple parts of the target person.
[0012] In some embodiments, in the person re-identification method, when extracting an overall feature vector of the image of the target person for the detail feature vector, and the overall feature vector represents the overall feature information of the target person, and the overall feature information includes the detail information of multiple parts of the target person, it includes:
[0013] Extract the position information of each part of the target person from the detailed feature vector, and add the position information of each part to the detailed feature vector to update the detailed feature vector. Each position information is correspondingly associated with the detailed information of the part corresponding to the position information, so that the updated detailed feature vector represents the overall feature information of the target person through the association relationship between the position information and the detailed information of each part, and the overall feature information contains the detailed information of multiple parts of the target person;
[0014] Extract a person feature vector from the updated detailed feature vector, which represents the overall feature information of the target person and the overall feature information contains the detailed information of multiple parts of the target person.
[0015] In some embodiments, for the person re-identification method, when determining whether there is a monitoring feature in each person feature library in the monitoring feature library whose similarity with the to-be-retrieved feature meets a first preset condition for the to-be-retrieved feature, the following judgment steps are included:
[0016] Calculate the vector similarity between the to-be-retrieved feature and the monitoring features in the person feature library respectively, and determine a preset number of monitoring features with the highest vector similarities;
[0017] Judge whether there is an alternative monitoring feature whose vector similarity with the to-be-retrieved feature meets a preset threshold among the determined preset number of monitoring features with the highest vector similarities.
[0018] In some embodiments, for the person re-identification method, the steps of calculating the vector similarity between the to-be-retrieved feature and the monitoring features in the person feature library respectively, and determining a preset number of monitoring features with the highest vector similarities include the following steps:
[0019] According to the shooting time information corresponding to the monitoring features in the person feature library, screen out the monitoring features in a preset time period, where the preset time period is a time period determined according to the input time signal, or a preset time period with a preset duration closest to the current moment;
[0020] Calculate the vector similarity between the to-be-retrieved feature and the screened monitoring features respectively, and determine a preset number of monitoring features with the highest vector similarities.
[0021] In some embodiments, in the person feature library of the person re-identification method, each monitoring feature also corresponds to a first label;
[0022] When determining the recognition result based on the monitoring feature whose similarity meets the first preset condition, the recognition result also includes the first label corresponding to the monitoring feature that meets the first preset condition;
[0023] After determining the recognition result based on the monitoring feature that meets the first preset condition, the method further includes:
[0024] According to the first label corresponding to the monitoring feature that meets the first preset condition, determine an image matching the first label from a pre-constructed monitoring image library.
[0025] Add the image matching the first label to the recognition result to update the recognition result.
[0026] In some embodiments, in the person re-identification method, the monitoring feature library is constructed by the following construction method, and the construction method includes:
[0027] Construct a person feature library corresponding to the camera, and add the attribute information of the camera to the person feature library;
[0028] Convert the video collected by the camera into an initial image, and determine whether there is a person in the initial image;
[0029] If there is, process the initial image with a person, and extract the person feature vector of the person image from the processed initial image as a monitoring feature, and store the monitoring feature and the shooting time of the initial image in the person feature library.
[0030] In some embodiments, in the person re-identification method, after determining the recognition result based on the monitoring feature that meets the first preset condition, the method further includes:
[0031] Determine the movement trajectory of the target person according to the attribute information of the camera in the recognition result.
[0032] In some embodiments, a person re-identification device is further provided, which is applied to a monitoring system. The monitoring system includes at least one camera and a pre-constructed monitoring feature library; the monitoring feature library includes at least one person feature library; each person feature library corresponds to a camera, and the person feature library includes the attribute information of the camera and the monitoring features of each person image captured by the camera, where the monitoring feature of each person image corresponds to the shooting time information of the person image; the device includes:
[0033] An acquisition module, configured to acquire an image of a target person, and extract the person feature vector of the target person from the image of the target person as a feature to be retrieved;
[0034] A first judgment module, configured to determine, for the to-be-retrieved feature, whether there is a monitoring feature in each person feature library in the monitoring feature library whose similarity to the to-be-retrieved feature meets a first preset condition;
[0035] A first determination module, configured to, if there is a monitoring feature in the monitoring feature library whose similarity to the to-be-retrieved feature meets the first preset condition, determine an identification result according to the monitoring feature whose similarity meets the first preset condition, where the identification result includes the shooting time information corresponding to the monitoring feature that meets the first preset condition, and the camera attribute information in the person feature library where the monitoring feature that meets the first preset condition exists.
[0036] In some embodiments, an electronic device is further provided, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus, and the processor executes the machine-readable instructions to perform the steps of the person re-identification method.
[0037] In the person re-identification method described in this application, the person images contained in the videos captured by each camera are pre-converted into person feature vectors, and the person feature vectors are used as monitoring features. The monitoring features are classified and stored according to the cameras. When it is necessary to determine the action trajectory of the target person, only the person feature vector of the target person needs to be extracted from the image of the target person in real time as the to-be-retrieved vector. Using the to-be-retrieved vector, traverse and retrieve whether there is a monitoring feature whose similarity meets the requirements in each person feature library, without the need to perform steps such as image interception, image preprocessing, and person feature extraction on the videos captured by the cameras in real time. Using vector retrieval technology increases the retrieval efficiency and improves the efficiency and real-time performance of finding the target person in the videos captured by the monitoring system. Description of the Drawings
[0038] To more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 Shows the method flowchart of the person re-identification method described in the embodiments of this application;
[0040] Figure 2 Shows the method flowchart for preprocessing the image of the target person in the embodiments of this application;
[0041] Figure 3 The figure shows a flowchart of a method for extracting a person feature vector of a target person according to an embodiment of the present application;
[0042] Figure 4 The figure shows obtaining the image I i of the original feature map F i schematic process diagram;
[0043] Figure 5 The figure shows the divided block feature map CF obtained according to an embodiment of the present application j schematic diagram;
[0044] Figure 6 The figure shows the block feature map CF according to an embodiment of the present application j and the original feature map F i schematic diagram of the position correspondence relationship;
[0045] Figure 7 The figure shows a flowchart of a method for constructing a monitoring feature library according to an embodiment of the present application;
[0046] Figure 8 The figure shows a schematic structural diagram of a person re-identification device according to an embodiment of the present application;
[0047] Figure 9 The figure shows a schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to actual scale. The flowcharts used in the present application show operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0049] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. The components of the embodiments of the present application usually described and illustrated in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0050] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the features stated thereafter, but does not exclude the addition of other features.
[0051] With the continuous development of society, at present in the market, video surveillance technologies and systems have been widely applied. For example, a certain number of cameras have been installed in office areas, squares, banks, streets, etc., realizing the functions of real-time collection, viewing, and storage of surveillance videos in each area. With the popularization of video surveillance, it has brought certain conveniences to all walks of life. For example: First, in the field of security, it has, to a certain extent, prevented the occurrence of bad behaviors (such as theft and robbery) and reduced the technical barriers for investigators to break through bad events (such as theft and robbery); Second, for places where personnel need to patrol, the responsible personnel can also confirm whether the patrol personnel have conducted regular inspections of specific equipment on time through video surveillance recordings. Third, for determining the movement trajectories of missing persons, the areas where the missing persons appear can be quickly located. However, for the situations described above that require finding specific persons from the cameras, it is necessary for the responsible personnel to check the video surveillance recordings one by one, which is a heavy task, not only time-consuming and laborious, but also inefficient.
[0052] With the continuous development of technology, some products that automatically find suspicious persons or the trajectories of specific patrol personnel from videos recorded by different cameras have also emerged on the market. However, the methods adopted by these products are all realized based on the matching between the personal characteristics of specific persons and the personal characteristics in the camera recordings. When the number of cameras is large, it takes a long time to extract the personal characteristics of the images in the camera recordings one by one. Therefore, the search efficiency is low, resulting in a waste of time. For the surveillance system, real-time retrieval cannot be achieved, which is not conducive to the popularization and use of this kind of product.
[0053] Based on this, the present application proposes a person re-identification method, which is applied to a monitoring system. The monitoring system includes at least one camera and a pre-constructed monitoring feature library; the monitoring feature library includes at least one person feature library; each person feature library corresponds to a camera, and the person feature library includes the attribute information of the camera and the monitoring features of each person image captured by the camera. Among them, the monitoring feature of each person image corresponds to the shooting time information of the person image; as Figure 1 shown, the method includes:
[0054] S101. Obtain an image of a target person, and extract a person feature vector of the target person from the image of the target person as a feature to be retrieved;
[0055] S102. For the feature to be retrieved, determine whether there is a monitoring feature in each person feature library in the monitoring feature library whose similarity to the feature to be retrieved meets a first preset condition;
[0056] S103. If there is, determine an identification result according to the monitoring feature whose similarity meets the first preset condition. The identification result includes the shooting time information corresponding to the monitoring feature that meets the first preset condition, and the camera attribute information in the person feature library where the monitoring feature that meets the first preset condition exists.
[0057] In the pre-constructed monitoring feature library, a person feature library is constructed for each camera respectively. Among them, the attribute information of the camera in the person feature library is at least one of the following: the number of the camera and the geographical location of the camera. For example: the attribute information of a certain camera is: at the door of Ward 1111, 3rd Floor, Building 5 of the hospital; the attribute information of a certain camera is: QQ Road Section, Camera No. 3.
[0058] In the embodiment of the present application, the monitoring feature of each person image captured by the camera is the person feature vector of the person image.
[0059] The pre-constructed monitoring feature library is stored in the storage module of the monitoring system. The storage module of the monitoring system can be a cloud, a cloud server, a server, etc.
[0060] The person re-identification method described in this application pre-converts the person images contained in the videos captured by each camera into person feature vectors, uses the person feature vectors as monitoring features, and stores the monitoring features classified by camera. When it is necessary to determine the action trajectory of a target person, it is only necessary to extract the person feature vector of the target person from the image of the target person in real time as the vector to be retrieved. Using the vector to be retrieved, traverse and search whether there are monitoring features with a similarity meeting the requirements in each person feature library, without the need to perform steps such as real-time image capture, image preprocessing, and person feature extraction on the videos captured by the cameras. Using vector retrieval technology increases the retrieval efficiency and improves the efficiency and real-time performance of finding the target person in the videos captured by the monitoring system.
[0061] Furthermore, the pre-constructed monitoring feature library uses the Milvus massive feature storage service to store each monitoring feature in each person feature library, overcoming the problem of slow speed of brute-force matching of massive vectors, so as to quickly implement the retrieval of monitoring features later.
[0062] S101. Obtain an image of the target person, and extract the person feature vector of the target person from the image of the target person as the feature to be retrieved.
[0063] Specifically, the image of the target person refers to an image that only contains a single target person, and the area of the person occupies more than 95% of the total area of the image. From the image of the target person, extract the normalized feature expression of the person image data, such as [0.001, 0.00256, 0,......, 0.003], a total of 256 dimensions, as the person feature vector.
[0064] That is to say, when the original image containing a single target person is obtained, it is necessary to preprocess the original image of the target person to obtain an image in which the area of the person occupies more than 95% of the total area of the image.
[0065] The preprocessing of the image of the target person is as Figure 2 shown, and includes the following steps:
[0066] S201. Scale the original image containing a single target person to a preset specification to obtain a scaled image;
[0067] S202. Judge whether the ratio between the number of pixels occupied by the target person in the scaled image and the total number of pixels occupied by the image is greater than a preset threshold.
[0068] S203. If it is greater than, determine that the scaled image is the image of the target person. If it is less than or equal to, repeat the above steps S201 and S202 until the ratio between the number of pixels occupied by the target person and the total number of pixels of the image is greater than the preset threshold.
[0069] In the embodiments of the present application, the preset threshold is [95%].
[0070] Specifically, in step S201, the original image containing a single target person is scaled to an image I of (64, 128, 3) dimensions. At the same time, calculate the total number of pixels IT of the scaled (64, 128, 3)-dimensional image.
[0071] In step S202, send the image I into a person segmenter to obtain the total number of pixels PT of the person in the image I; calculate the percentage of the ratio of PT to IT, and determine whether the percentage is greater than 95%.
[0072] Preferably, the person segmenter is Mask-Rcnn.
[0073] If it is greater than, determine that the scaled image I is the image of the target person.
[0074] When there are n original images containing a single target person, preprocess the original images respectively, and set the preprocessed images as image I i , i ∈ [1, n], and store them in the queue Qimg, Qimg = {I1, I2,..., I i ,..., I n}.
[0075] Currently, products that search for suspicious persons or the trajectories of specific patrol personnel from videos recorded by different cameras adopt a single convolutional neural network cnn (convolution neural network) technology to extract human features, and perform a search for specific persons through a brute-force matching method. The disadvantages of this method are: due to the inductive bias of cnn, the extracted human features are relatively one-sided, and it is easy to cause a large number of misidentifications and missed identifications.
[0076] In the embodiments of the present application, extract the human feature vector of the target person from the image of the target person as the feature to be retrieved. As Figure 3 shown, the extraction method includes:
[0077] S301. For the image of the target person, extract the detailed feature vector of the image of the target person, and the detailed feature vector represents the detailed information of multiple parts of the target person;
[0078] S302. Extract the overall feature vector of the image of the target person for the detailed feature vector. The overall feature vector represents the overall feature information of the target person, and the overall feature information includes the detailed information of multiple parts of the target person.
[0079] For the multi-camera scenario without overlapping regions, where the clothing, posture, and appearance of a single person hardly change in a short period of time, first extract the detailed features of the target person, such as facial organ features (the shape and curvature of the eyes, the curvature of the eyebrows), etc. Then, further consider all aspects to obtain the overall features (such as the clothing color of the target person, the height, weight, and build of the body), overcoming the problem that the person features extracted by a single extraction method are relatively one-sided, making the extracted person feature vector more comprehensive and capable of expressing the features of the target person more accurately, thereby improving the accuracy of the subsequent target person retrieval results.
[0080] In S302, for the detailed feature vector, extract the overall feature vector of the image of the target person. The overall feature vector represents the overall feature information of the target person, and the overall feature information includes the detailed information of multiple parts of the target person. This can be done in various ways. For example: first extract the feature vector that only represents the overall feature information from the target person graph composed of the detailed feature vector, and then fuse the remaining detailed feature vectors to obtain the overall feature vector that represents the overall feature information of the target person and includes the detailed information of multiple parts of the target person in the overall feature information.
[0081] In the embodiment of the present application, specifically, for the detailed feature vector, extract the overall feature vector of the image of the target person. The overall feature vector represents the overall feature information of the target person, and the overall feature information includes the detailed information of multiple parts of the target person, including:
[0082] S3021. Extract the position information of each part of the target person from the detailed feature vector, and add the position information of each part to the detailed feature vector to update the detailed feature vector. Each position information is associated with the detailed information of the part corresponding to the position information, so that the updated detailed feature vector represents the overall feature information of the target person through the association relationship between the position information and the detailed information of each part, and the overall feature information includes the detailed information of multiple parts of the target person;
[0083] S3022. Extract from the updated detailed feature vector the person feature vector that represents the overall feature information of the target person and includes the detailed information of multiple parts of the target person.
[0084] In this extraction method, position information is added to the original detailed feature vector, so that the updated detailed feature vector can represent the overall feature information while still retaining the detailed features of multiple parts of the target person, thereby making the process of extracting the overall feature vector more concise and the extraction efficiency higher.
[0085] Specifically, a pre-constructed extraction model is used to extract the person feature vector of the target person.
[0086] The pre-constructed extraction model includes a detailed feature extraction module and an overall feature extraction module;
[0087] The detailed feature extraction module is used to extract the detailed feature vector of the image of the target person for the image of the target person. The detailed feature vector represents the detailed information of multiple parts of the target person, and sends the detailed feature vector to the overall feature extraction module;
[0088] The overall feature extraction module is used to receive the detailed feature vector, extract the position information of each part of the target person from the detailed feature vector, and add the position information of each part to the detailed feature vector to update the detailed feature vector. Each position information is associated with the detailed information of the part corresponding to the position information, so that the updated detailed feature vector represents the overall feature information of the target person through the association relationship between the position information and the detailed information of each part, and the overall feature information includes the detailed information of multiple parts of the target person; from the updated detailed feature vector, extract the person feature vector representing the overall feature of the target person.
[0089] Specifically, the detailed feature extraction module uses a CNN convolutional neural network, and the overall feature extraction module uses a transformers network.
[0090] For the characteristics that in a multi-camera scenario without overlapping areas, the clothing, posture, and appearance of a single person hardly change in a short period of time, a CNN convolutional neural network focusing on texture information is used to extract the detailed features of the target person, such as facial organ features (the shape of the eyes, the curvature, the curvature of the eyebrows), etc. Then, based on the extracted detailed feature vector, a transformers network based on the self-attention mechanism is used to further consider all aspects to obtain the overall features (such as the clothing color of the target person, the height, weight, and build of the body), overcoming the inductive bias of the CNN, making the extracted person feature vector more comprehensive and able to express the features of the target person more accurately, thereby improving the accuracy of the subsequent target person retrieval results.
[0091] When extracting the detailed feature vector, high requirements are imposed on the quality of image details, while when extracting the overall features, the requirements for the quality of image details are relatively low. Therefore, first use the cnn convolutional neural network to extract the detailed feature vector of the target person's image, and then use the transformers network to extract the overall features.
[0092] For the queue Qimg, Qimg = {I1, I2,..., I i ,..., I n}, for each image I i , i ∈ [1, n], the extraction method specifically includes the following steps:
[0093] S401. Send each image I img in the queue Q i , i ∈ [1, n] into the convolutional neural network with ResNet50 as the backbone network to obtain the original feature map F i of the image I i . The dimension of the original feature map F i is: 1 * 1024; the specific process is as Figure 4 shown:
[0094] S402. Divide the original feature map F i into 128 1 * 8 block feature maps with a length of 8 to meet the input of the transformers structure and mark the position of each block of features at the same time;
[0095] As Figure 5 shown, the obtained block feature map after division is denoted as CF j . The dimension of each block feature map CF j is 1 * 8, where j ∈ [1, 128];
[0096] S403. Record the starting position and ending position of each block feature map CF j in the original feature map F i , and use the starting position xi and the ending position xi + 7 to mark the position X j of the block feature map CF i in the original feature map F j , where X j = (xi, xi + 7) is the position encoding of the block feature map CF j ; the position correspondence between the block feature map CF j and the original feature map F i is as Figure 6 shown;
[0097] S404. Connect the position encoding X j and the block feature map CF j, obtain the input vector Z = {Z j}, where Z j = concat(X j , CF j ), where j ∈ [1, 128];
[0098] S405. Input the input vector Z into the transformers network to calculate the feature FE of the target person, where FE = T(Z1, Z2,..., Z 128 ), and finally obtain the image feature expression of the target person, denoted as FE, where FE = {FE1, FE2, FE3,..., Fv,..., FE 256}, with a total of 256 dimensions, where v ∈ [1, 256]; Fv is a one-dimensional vector in FE;
[0099] S406. Normalize the image feature expression FE vector of the target person to obtain a normalized feature vector as the person feature vector of the target person.
[0100] Through the above steps, the extraction of the person feature vector of the target person image is completed.
[0101] In the embodiment of the present application, the monitoring feature library is constructed by the following construction method, as Figure 7 shown, and the construction method includes:
[0102] S701. Construct a person feature library corresponding to the camera and add the attribute information of the camera to the person feature library;
[0103] S702. Convert the video collected by the camera into an initial image and determine whether there is a person in the initial image;
[0104] S703. If there is, process the initial image with a person and extract the person feature vector of the person image from the processed initial image as a monitoring feature, and store the monitoring feature and the shooting time of the initial image in the person feature library.
[0105] In step S802, through the decoding server, the video collected by the camera is decoded into an image using ffmpeg, and the obtained initial image after decoding is sent to the person detector.
[0106] The person detector detects whether there is a person in the initial image and returns the coordinates of the person in the initial image when there is a person. The coordinates of the person include (x, y, w, h), where x and y represent the upper left x-coordinate and upper left y-coordinate of the person in the image respectively, w represents the width of the person in the image, and h represents the height of the person in the image. Preferably, the person detector is PP-Yolo.
[0107] The person feature extractor extracts person data from the initial image according to the coordinates of the person returned by the person detector. After scaling this person data into data with a dimension of (64, 128, 3), the scaled data is sent into the extraction model to extract person features, obtaining a normalized feature representation of the person image data, such as [0.001, 0.00256, 0,..., 0.003], with a total of 256 dimensions.
[0108] The person data is scaled into data with a dimension of (64, 128, 3), and the scaling process is as described in steps S201, S202, and S203 in this embodiment.
[0109] The step of sending the scaled data into the extraction model to extract person features and obtaining a normalized feature representation of the person image data, such as [0.001, 0.00256, 0,..., 0.003], with a total of 256 dimensions, is as described in S301 and S302 in this embodiment, and specifically as described in steps S401 - S406 in this embodiment.
[0110] In the embodiment of the present application, when constructing the person feature library, the method of extracting the monitoring feature from the initial image containing a person is the same as the method of extracting the vector to be retrieved from the image of the target person. Both use an extraction model composed of a convolutional neural network and a transformers network for extraction, so that for the same person, the extracted monitoring feature and the feature to be retrieved are more similar, thereby improving the retrieval accuracy.
[0111] In the embodiment of the present application, for the feature to be retrieved, determining whether there is a monitoring feature in each person feature library in the monitoring feature library whose similarity with the feature to be retrieved satisfies the first preset condition includes the following judgment steps:
[0112] Calculate the vector similarity between the feature to be retrieved and the monitoring features in the person feature library respectively, and determine a preset number of monitoring features with the highest vector similarities;
[0113] Judge whether there is an alternative monitoring feature among the determined preset number of monitoring features with the highest vector similarities whose vector similarity with the feature to be retrieved satisfies the preset threshold.
[0114] In the embodiment of the present application, the specific judgment steps are as follows:
[0115] S407: Use the person feature vector of the target person obtained in step S406 as the person feature to be retrieved, denoted as TF, where TF is 256-dimensional;
[0116] S408: Obtain all cameras in the monitoring system, and denote the Kth camera as CAM k , where k ∈ [1, p], and p is the total number of cameras in the monitoring system;
[0117] S409: Obtain the person feature vectors in the person feature library corresponding to the camera CAM k , denoted as LF kg , where k ∈ [1, p], p is the total number of cameras in the monitoring system, g ∈ [1, m(k)], m(k) is the total number of all monitoring features in the person feature library corresponding to the kth camera, and the dimension of LF kg is 256;
[0118] S410: Use the high-speed vector retrieval service of Milvus to calculate the inner product distance between TF and each LF kg , and store the inner product distance into the queue Q dis ={Dis(TF, LF k1 ), Dis(TF, LF k2 ), Dis(TF, LF k3 ),..., Dis(TF, LF km(k) )};
[0119] S411: Sort the inner product distances in the Q dis queue in descending order, and take out the first 5 inner product distances after the descending order, denoted as DisDesc = {DT1, DT2, DT3, DT4, DT5};
[0120] S412: If DT c > 0.55, where c ∈ {1, 2, 3, 4, 5}, it means that the target appears in the current camera, and store the shooting time information corresponding to DT c and the attribute information of this camera into the queue ReCamList; if the inner product distances in DisDesc are all less than 0.55, it means that the person to be searched does not appear in the current camera;
[0121] S413: Repeat steps S409 to S412 until all person feature libraries corresponding to the cameras are searched;
[0122] S414. Obtain the attribute information of the camera where the target person appears and the shooting time information when the target person appears from the queue ReCamList.
[0123] Through two screenings, the first screening selects a preset number of monitoring features with the highest vector similarity, and the second only needs to calculate whether these preset number of monitoring features with the highest vector similarity meet the preset threshold, reducing the comparison calculation between the vector similarity and the preset threshold, and increasing the retrieval speed and efficiency.
[0124] To further reduce the calculation amount between calculating the vector similarity of the vector to be retrieved and the vectors in the person feature library, calculate the vector similarity between the feature to be retrieved and the monitoring features in the person feature library respectively, and determine a preset number of monitoring features with the highest vector similarity; the method includes the following steps:
[0125] According to the shooting time information corresponding to the monitoring features in the person feature library, screen out the monitoring features in the preset time period, where the preset time period is the time period determined according to the input time signal, or the preset time period with a preset duration closest to the current moment;
[0126] Calculate the vector similarity between the feature to be retrieved and the screened monitoring features respectively, and determine a preset number of monitoring features with the highest vector similarity.
[0127] That is to say, only the similarity between the monitoring features in the preset time period and the vector to be retrieved needs to be calculated, greatly reducing the calculation amount when calculating the vector similarity.
[0128] The preset time period is the time period determined according to the input time signal, and the input time signal is determined according to the recognition purpose of the target person. For example, to determine whether the patrol personnel conduct patrol according to the preset route between 7 pm and 8 pm on the 8th, only the similarity between the monitoring feature vector with the corresponding shooting time information between 7 pm and 8 pm on the 8th and the person feature vector of the patrol personnel needs to be calculated.
[0129] In the embodiment of the present application, in the described person re-identification method, in the person feature library, each monitoring feature also corresponds to a first label;
[0130] When determining the recognition result according to the monitoring features whose similarity meets the first preset condition, the recognition result also includes the first label corresponding to the monitoring features that meet the first preset condition;
[0131] After determining the recognition result according to the monitoring features whose similarity meets the first preset condition, the method further includes:
[0132] Based on the first label corresponding to the monitoring feature that meets the first preset condition, determine the image that matches the first label from the pre-constructed monitoring image library.
[0133] Add the image that matches the first label to the recognition result to update the recognition result.
[0134] The monitoring image library includes sub-image libraries corresponding one-to-one to the person feature libraries. In each sub-image library, the initial images decoded from the video stream of the camera and corresponding one-to-one to the monitoring features are stored, and the initial images include the person corresponding to the monitoring feature.
[0135] When the camera information and shooting time information of the target person are captured, the image of the target person captured by the camera is also output at the same time, which more intuitively displays the retrieval result and can also be used as evidence and proof.
[0136] The described monitoring image library can be stored locally (such as on a hard disk, etc.), or can be stored on a server, in the cloud, etc.
[0137] The first label can be the camera number of the monitoring feature + the monitoring feature serial number. In the name of the corresponding image in the sub-image library in the monitoring image library, this first label (camera number + monitoring feature serial number) is included, so as to find the image whose name contains the first label in the corresponding sub-image library in the monitoring image library.
[0138] After determining the recognition result according to the monitoring feature whose similarity meets the first preset condition, the method further includes:
[0139] Determine the movement trajectory of the target person according to the attribute information of the camera in the recognition result.
[0140] In some embodiments, the position information in the attribute information of the camera can be displayed on an electronic map, where each position information corresponds to a position identifier on the electronic map, so as to more intuitively display the movement trajectory of the target person.
[0141] Such as Figure 8 As shown, an embodiment of the present application further provides a person re-identification device applied to a monitoring system. The monitoring system includes at least one camera and a pre-constructed monitoring feature library; the monitoring feature library includes at least one person feature library; each person feature library corresponds to a camera, and the person feature library includes the attribute information of the camera and the monitoring features of each person image captured by the camera, where the monitoring feature of each person image corresponds to the shooting time information of the person image; the device includes:
[0142] An acquisition module 801, configured to acquire an image of a target person, and extract a person feature vector of the target person from the image of the target person as a feature to be retrieved;
[0143] A first judgment module 802, configured to, for the feature to be retrieved, judge whether there is a monitoring feature in each person feature library in the monitoring feature library whose similarity with the feature to be retrieved meets a first preset condition;
[0144] A first determination module 803, configured to, if there is a monitoring feature in the monitoring feature library whose similarity with the feature to be retrieved meets the first preset condition, determine an identification result according to the monitoring feature whose similarity meets the first preset condition, where the identification result includes shooting time information corresponding to the monitoring feature that meets the first preset condition, and camera attribute information in the person feature library where the monitoring feature that meets the first preset condition exists.
[0145] In the person re-identification device according to an embodiment of the present application, the acquisition module 801 further includes:
[0146] A first extraction module, configured to, for an image of a target person, extract a detail feature vector of the image of the target person, where the detail feature vector represents detail information of multiple parts of the target person;
[0147] A second extraction module, configured to, for the detail feature vector, extract an overall feature vector of the image of the target person, where the overall feature vector represents overall feature information of the target person, and the overall feature information includes detail information of multiple parts of the target person.
[0148] The second extraction module is specifically configured to extract position information of each part of the target person from the detail feature vector, and add the position information of each part to the detail feature vector to update the detail feature vector, and each position information is correspondingly associated with the detail information of the part corresponding to the position information, so that the updated detail feature vector represents the overall feature information of the target person through the association relationship between the position information and the detail information of each part, and the overall feature information includes detail information of multiple parts of the target person; extract a person feature vector from the updated detail feature vector, where the person feature vector represents the overall feature information of the target person and the overall feature information includes detail information of multiple parts of the target person.
[0149] The first judgment module 802 includes:
[0150] A calculation module, configured to calculate vector similarities between the feature to be retrieved and monitoring features in a person feature library respectively, and determine a preset number of monitoring features with the highest vector similarities;
[0151] A second judgment module, configured to judge whether there is an alternative monitoring feature among the determined monitoring features with the highest vector similarity of the preset number, and the vector similarity between the alternative monitoring feature and the feature to be retrieved meets a preset threshold.
[0152] The calculation module is specifically configured to: screen out monitoring features in a preset time period according to the shooting time information corresponding to the monitoring features in the person feature library, where the preset time period is a time period determined according to the input time signal, or a preset time period with a preset duration closest to the current moment; calculate the vector similarity between the feature to be retrieved and the screened monitoring features respectively, and determine a preset number of monitoring features with the highest vector similarity.
[0153] In an embodiment of the present application, in the person re-identification device, in the person feature library, each monitoring feature also corresponds to a first mark; when determining the recognition result according to the monitoring feature whose similarity meets the first preset condition, the recognition result also includes the first mark corresponding to the monitoring feature that meets the first preset condition; the device further includes:
[0154] A second determination module, configured to, after determining the recognition result according to the monitoring feature whose similarity meets the first preset condition, determine an image matching the first mark from a pre-constructed monitoring image library according to the first mark corresponding to the monitoring feature that meets the first preset condition;
[0155] An adding module, configured to add the image matching the first mark to the recognition result to update the recognition result.
[0156] In an embodiment of the present application, the person re-identification device further includes:
[0157] A construction module, configured to construct the monitoring feature library.
[0158] The construction module is specifically configured to construct a person feature library corresponding to a camera, and add the attribute information of the camera to the person feature library;
[0159] Convert the video collected by the camera into an initial image, and judge whether there is a person in the initial image;
[0160] If there is, process the initial image with a person, extract the person feature vector of the person image from the processed initial image as a monitoring feature, and store the monitoring feature and the shooting time of the initial image in the person feature library.
[0161] The person re-identification method device described in the embodiment of the present application further includes:
[0162] A third determination module, configured to determine the movement trajectory of the target person according to the attribute information of the camera in the recognition result after determining the recognition result according to the monitoring features that satisfy the first preset condition.
[0163] As Figure 9 shown, an embodiment of the present application further provides an electronic device, including: a processor 901, a memory 902, and a bus 903. The memory 902 stores machine-readable instructions executable by the processor 901. When the electronic device runs, the processor 901 communicates with the memory 902 through the bus 903, and the processor 901 executes the machine-readable instructions to perform the steps of the person re-identification method.
[0164] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the person re-identification method are executed.
[0165] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the method embodiments, which will not be elaborated herein. In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.
[0166] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0167] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0168] When the above-described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a platform server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0169] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A person re-identification method, characterized in that, Applied to a monitoring system, the monitoring system includes at least one camera and a pre - constructed monitoring feature library; the monitoring feature library includes at least one person feature library; each person feature library corresponds to a camera, and the person feature library includes the attribute information of the camera and the monitoring features of each person image captured by the camera, wherein the monitoring feature of each person image corresponds to the shooting time information of that person image; the method includes: Obtain an image of a target person, and extract the person feature vector of the target person from the image of the target person as the feature to be retrieved; For the feature to be retrieved, determine whether there is a monitoring feature in each person feature library in the monitoring feature library whose similarity with the feature to be retrieved meets a first preset condition; If there is, determine the recognition result according to the monitoring feature whose similarity meets the first preset condition. The recognition result includes the shooting time information corresponding to the monitoring feature that meets the first preset condition, and the camera attribute information in the person feature library where the monitoring feature that meets the first preset condition exists; Extracting the person feature vector of the target person from the image of the target person as the feature to be retrieved includes: Put the queue Q img Each image I i , i ∈ [1, n] into a convolutional neural network with ResNet50 as the backbone network to obtain the original feature map F i of the image I i , the original feature map F i The dimension is: 1 * 1024; The original feature map F i is divided into 128 1*8 block feature maps with a length of 8 to meet the input of the transformers structure while marking the position of each block of features; the obtained block feature map is denoted as CF j , and each block feature map CF j has a dimension of 1*8, where j ∈ [1, 128]; Record each block feature map CF j at the starting position and ending position in the original feature map F i and mark the block feature map CF with the starting position xi and the ending position xi+7 j at the position X i in the original feature map F j where X j =(xi, xi+7) is the position encoding of the block feature map CF j ; Connection position encoding X j and the block feature map CF j to obtain the input vector Z = {Z j}, where Z j = concat(X j , CF j ), where j ∈ [1, 128]; Input the input vector Z into the transformers network to calculate the features FE of the target person, where FE = T(Z1, Z2,..., Z 128 ), and finally obtain the image feature representation of the target person, denoted as FE, where FE = {FE1, FE2, FE3,..., Fv,..., FE 256}, with a total of 256 dimensions, where v ∈ [1, 256]; Fv is a one-dimensional vector in FE; Normalize the image feature expression FE vector of the target person to obtain a normalized feature vector as the person feature vector of the target person.
2. The person re-identification method according to claim 1, wherein The extracting the person feature vector of the target person from the image of the target person as the feature to be retrieved includes: For the image of the target person, extract the detailed feature vector of the image of the target person, and the detailed feature vector represents the detailed information of multiple parts of the target person; For the detailed feature vector, extract the overall feature vector of the image of the target person, and the overall feature vector represents the overall feature information of the target person, and the overall feature information contains the detailed information of multiple parts of the target person.
3. The person re - identification method according to claim 2, wherein For the detailed feature vector, extracting the overall feature vector of the image of the target person, and the overall feature vector represents the overall feature information of the target person, and the overall feature information contains the detailed information of multiple parts of the target person, includes: Extract the position information of each part of the target person from the detailed feature vector, and add the position information of each part to the detailed feature vector to update the detailed feature vector. Each position information is associated with the detailed information of the part corresponding to the position information, so that the updated detailed feature vector represents the overall feature information of the target person through the association relationship between the position information and the detailed information of each part, and the overall feature information contains the detailed information of multiple parts of the target person; Extract from the updated detailed feature vector a person feature vector that represents the overall feature information of the target person and the overall feature information contains the detailed information of multiple parts of the target person.
4. The person re-identification method according to claim 1, wherein For the to-be-retrieved feature, determining whether there is a monitoring feature in each person feature library in the monitoring feature library whose similarity to the to-be-retrieved feature meets a first preset condition includes the following determination steps: Calculate the vector similarity between the to-be-retrieved feature and the monitoring features in the person feature library respectively, and determine a preset number of monitoring features with the highest vector similarities; Judge whether there is an alternative monitoring feature among the determined preset number of monitoring features with the highest vector similarities whose vector similarity to the to-be-retrieved feature meets a preset threshold.
5. The person re-identification method according to claim 4, wherein Calculate the vector similarity between the to-be-retrieved feature and the monitoring features in the person feature library respectively, and determine a preset number of monitoring features with the highest vector similarities; including the following steps: According to the shooting time information corresponding to the monitoring features in the person feature library, filter out the monitoring features in a preset time period, where the preset time period is a time period determined according to the input time signal, or a preset time period with a preset duration closest to the current moment; Calculate the vector similarity between the to-be-retrieved feature and the filtered monitoring features respectively, and determine a preset number of monitoring features with the highest vector similarities.
6. The person re-identification method according to claim 1, wherein In the person feature library, each monitoring feature also corresponds to a first label; When determining the recognition result according to the monitoring feature whose similarity meets the first preset condition, the recognition result also includes the first label corresponding to the monitoring feature that meets the first preset condition; After determining the recognition result according to the monitoring feature whose similarity meets the first preset condition, the method further includes: According to the first label corresponding to the monitoring feature that meets the first preset condition, determine the image matching the first label from the pre-constructed monitoring image library; Add the image matching the first label to the recognition result to update the recognition result.
7. The person re-identification method according to claim 1, wherein The monitoring feature library is constructed by the following construction method, and the construction method includes: Construct a person feature library corresponding to the camera, and add the attribute information of the camera to the person feature library; Convert the video collected by the camera into an initial image, and judge whether there is a person in the initial image; If so, process the initial image with a person, and extract the person feature vector of the person image from the processed initial image as the monitoring feature, and store the monitoring feature and the shooting time of the initial image in the person feature library.
8. The person re-identification method according to claim 1, wherein After determining the recognition result according to the monitoring feature whose similarity meets the first preset condition, the method further includes: According to the attribute information of the camera in the recognition result, determine the movement track of the target person.
9. A person re-identification device, characterized in that, Applied to a monitoring system, the monitoring system includes at least one camera and a pre-built monitoring feature library; the monitoring feature library includes at least one person feature library; each person feature library corresponds to a camera, and the person feature library includes the attribute information of the camera and the monitoring features of each person image captured by the camera, wherein the monitoring feature of each person image corresponds to the shooting time information of the person image; the device includes: An acquisition module, configured to acquire an image of a target person, and extract a person feature vector of the target person from the image of the target person as a feature to be retrieved; A first judgment module, configured to, for the feature to be retrieved, judge whether there is a monitoring feature in each person feature library in the monitoring feature library whose similarity with the feature to be retrieved satisfies a first preset condition; A first determination module, configured to, if there is a monitoring feature in the monitoring feature library whose similarity with the feature to be retrieved satisfies the first preset condition, determine an identification result according to the monitoring feature whose similarity satisfies the first preset condition, where the identification result includes the shooting time information corresponding to the monitoring feature that satisfies the first preset condition, and the camera attribute information in the person feature library where the monitoring feature that satisfies the first preset condition exists; Extracting a person feature vector of the target person from the image of the target person as a feature to be retrieved includes: Put the queue Q img Each image I i , i ∈ [1, n] into the convolutional neural network with ResNet50 as the backbone network, and obtain the original feature map F i of the image I i , the original feature map F i The dimension is: 1 * 1024; The original feature map F i is divided into 128 1*8 block feature maps with a length of 8 to meet the input of the transformers structure while marking the position of each block of features; the block feature maps obtained after division are denoted as CF j , and each block feature map CF j has a dimension of 1*8, where j ∈ [1, 128]; Record each block feature map CF j at the starting position and the ending position in the original feature map F i and mark the block feature map CF with the starting position xi and the ending position xi+7 j at the position X i in the original feature map F j where X j = (xi, xi+7) is the position encoding of the block feature map CF j ; Connection position encoding X j and the block feature map CF j to obtain the input vector Z = {Z j}, where Z j = concat(X j , CF j ), where j ∈ [1, 128]; Input the input vector Z into the transformers network to calculate the features FE of the target person, where FE = T(Z1, Z2,..., Z 128 ), and finally obtain the image feature expression of the target person, denoted as FE, where FE = {FE1, FE2, FE3,..., Fv,..., FE 256}, with a total of 256 dimensions, where v ∈ [1, 256]; Fv is a one-dimensional vector in FE; Normalizing the image feature expression FE vector of the target person to obtain a normalized feature vector as the person feature vector of the target person.
10. An electronic device, characterized in that, Including: A processor, a memory and a bus, the memory stores machine-readable instructions executable by the processor, when the electronic device runs, the processor communicates with the memory through the bus, and the processor executes the machine-readable instructions to perform the steps of the person re-identification method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Character trajectory retrieval method and system and computer readable storage medium
CN110532432A