A cross-modal person re-identification method

By training and converting the pedestrian recognition model on the HiSilicon Hi3516 platform, the problems of insufficient application of the existing mid-span modal pedestrian recognition in real environments and low algorithm efficiency are solved, and efficient pedestrian recognition is achieved through day and night linkage.

CN114764921BActive Publication Date: 2025-05-16SUN YAT SEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210421430.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-05-16
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

The existing cross-modal pedestrian re-identification technology is difficult to achieve day and night linkage in a real environment, and the algorithm has high power and large volume on the GPU, low computing efficiency on the CPU, and difficult to output results in real time.

Method used

By training the initial pedestrian recognition model and converting it into a cross-modal pedestrian recognition model, it is transplanted to the HiSilicon Hi3516 platform, combining powerful image processing capabilities and high computing performance, real-time image acquisition and pedestrian recognition are achieved.

Benefits of technology

The pedestrian re-identification is realized through day and night linkage, reducing CPU occupancy, accelerating the recognition speed, while maintaining the characteristics of low power consumption and small size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114764921B_ABST
    Figure CN114764921B_ABST
Patent Text Reader

Abstract

The present invention provides a cross-modal pedestrian re-identification method, which mainly includes the following steps: obtaining an initial pedestrian re-identification model through training with a first data set; converting the initial pedestrian re-identification model to obtain a cross-modal pedestrian re-identification model, and transplanting the cross-modal pedestrian re-identification model to a target platform; performing real-time image acquisition through the target platform, and extracting pedestrian detection results from the real-time images; inputting the pedestrian detection results and preset targets into the cross-modal pedestrian re-identification model to obtain pedestrian re-identification results, which can improve the real-world full-time and space application scenarios of pedestrian re-identification, establish daytime and nighttime joint pedestrian re-identification to realize offline or online cross-modal pedestrian re-identification, and the method can reduce CPU occupancy, speed up the cross-modal pedestrian re-identification speed, while maintaining the characteristics of low power consumption and small size, and can be widely used in the field of pedestrian re-identification technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian re-identification, and in particular to a cross-modal pedestrian re-identification method. Background Art

[0002] The pedestrian re-identification algorithm is a technology for matching detected pedestrians across cameras. It is derived from a subtask in Multi-Target Multi-Camera Tracking (MTMC). Pedestrian re-identification aims to use computer vision technology to match pictures or video sequences of the same pedestrian collected by different cameras. It has broad application prospects in the fields of intelligent security and smart transportation. For example, it can be combined with pedestrian detection, face recognition and other technologies to track suspects, track lost pedestrians, etc., to provide protection for people's lives and property.

[0003] In the relevant technical solutions, most surveillance cameras can automatically switch from RGB mode to infrared mode in the dark, that is, most surveillance cameras take RGB images during the day and infrared images at night. Since pedestrian re-identification generally uses image data taken by surveillance cameras to retrieve target people, the existence of huge modal differences between the above two types of image data has severed the spatiotemporal application scenarios of pedestrian re-identification in real environments, making it possible to use only RGB modality pedestrian re-identification for daytime queries and infrared modality pedestrian re-identification for nighttime queries. In addition, if the cross-modal pedestrian re-identification algorithm runs on a GPU, the high power consumption and large size of the GPU will make it inconvenient to actually lay out; if the cross-modal pedestrian re-identification algorithm runs on a CPU, the CPU's computing efficiency is low and it is difficult to output results in real time. Summary of the invention

[0004] In view of this, in order to at least partially solve the above technical problems, an object of an embodiment of the present invention is to provide a cross-modal pedestrian re-identification method that can better restore the real environment and achieve daytime and nighttime linkage.

[0005] To this end, the technical solution of the present application provides a cross-modal pedestrian re-identification method, comprising the following steps:

[0006] An initial person re-identification model is obtained by training the first data set;

[0007] Converting the initial person re-identification model to obtain a cross-modal person re-identification model, and transplanting the cross-modal person re-identification model to a target platform;

[0008] Real-time image acquisition is performed through the target platform, and pedestrian detection results are extracted from the real-time image;

[0009] The pedestrian detection result and the preset target are input into the cross-modal pedestrian re-identification model to obtain the pedestrian re-identification result.

[0010] In a feasible embodiment of the solution of the present application, the step of obtaining an initial person re-identification model through training with the first data set includes:

[0011] Acquire an RGB image and an infrared image from the first data set;

[0012] Performing feature extraction on the RGB image to obtain RGB modal features, and performing feature extraction on the infrared image to obtain infrared modal features;

[0013] The RGB modal feature and the infrared modal feature are combined to obtain a mixed modal feature;

[0014] A cross-modal feature is obtained according to the mixed modal feature output.

[0015] In a feasible embodiment of the solution of the present application, the step of obtaining the cross-modal feature according to the mixed modal feature output includes:

[0016] A cross-modal feature is obtained by outputting the mixed modal feature through a pooling layer; the cross-modal feature includes: identity-related features and identity-independent features.

[0017] In a feasible embodiment of the solution of the present application, the step of obtaining identity-related features and identity-independent features through the output of the pooling layer according to the mixed modal features includes:

[0018] Obtaining a first feature distance of two images that do not belong to the same class and a second feature distance of two images that belong to the same class in the mixed modal feature;

[0019] The identity-independent feature is determined by an inverse triplet loss function according to the first feature distance and the second feature distance.

[0020] In a feasible embodiment of the solution of the present application, the step of obtaining identity-related features and identity-independent features through a pooling layer output according to the mixed modal features includes:

[0021] The mixed modal features are extracted from the mixed modal features through correlation coefficient collaboration and center clustering loss function, triplet loss function and cross entropy loss function.

[0022] In a feasible embodiment of the solution of the present application, after the step of obtaining the cross-modal feature through the output of the pooling layer according to the mixed modal feature, the method further includes the following steps:

[0023] orthogonally decoupling the first eigenvector of the identity-related feature and the second eigenvector of the identity-independent feature through an orthogonal loss function;

[0024] The identity-irrelevant information is separated from the identity-relevant features according to the result of orthogonal decoupling.

[0025] In a feasible embodiment of the present application, the reverse triplet loss function is:

[0026] Loss Tri-reverse =max(d a,n -d a,p +α,0)

[0027] Among them, Loss Tri-reverse is the reverse triplet loss value, d a,n is the first feature distance between two images that do not belong to the same class, d a,p is the feature distance between two images belonging to the same category, and α is the reverse triplet loss value Loss Tri-reverse The boundary constant.

[0028] In a feasible embodiment of the present application, the correlation coefficient collaboration and center clustering loss function is:

[0029]

[0030] Among them, λ 1 and λ 2 is a hyperparameter, c is the number of classes in a mini-batch, is the RGB feature center vector of the i-th image in the small batch, is the infrared feature center vector of the i-th image in the small batch; is the relationship coefficient matrix of the i-th RGB feature in a small batch, is the relationship coefficient matrix of the i-th type of infrared features in a small batch, i = 1, 2, 3, …, n, and n is a positive integer.

[0031] In a feasible embodiment of the present application, the orthogonal loss function is:

[0032]

[0033] Where b. is the number of mini-batch inputs, is the identity-independent feature of the i-th image in the mini-batch input and Identity-related features for the i-th image in the input mini-batch.

[0034] In a feasible embodiment of the solution of the present application, the method converts the initial pedestrian re-identification model to obtain a cross-modal pedestrian re-identification model, and transplants the cross-modal pedestrian re-identification model to the target platform, including:

[0035] Performing a model framework verification on the initial person re-identification model;

[0036] It is determined that the initial person re-identification model does not belong to the Caffe framework, and the initial person re-identification model is converted into the cross-modal person re-identification model under the Caffe framework.

[0037] The advantages and beneficial effects of the present invention will be partially given in the following description, and the other parts can be understood through the specific embodiments of the present invention:

[0038] The technical solution of the present application can be based on the powerful image processing capabilities and high computing performance that the platform can provide, combined with a cross-modal pedestrian re-identification algorithm, to retrieve pedestrian information from video images in real time and extract feature values. It can improve the real-world full-time and space application scenarios of pedestrian re-identification, establish daytime and nighttime joint pedestrian re-identification to realize offline or online cross-modal pedestrian re-identification, and the method can reduce CPU occupancy and speed up cross-modal pedestrian re-identification, while maintaining the characteristics of low power consumption and small size. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0040] Figure 1 A flowchart of the steps of a cross-modal pedestrian re-identification method provided by the technical solution of this application;

[0041] Figure 2 This is a schematic diagram of the structure of the cross-modal pedestrian re-identification model in the technical solution of this application;

[0042] Figure 3 A flowchart of the steps of another cross-modal pedestrian re-identification method provided by the technical solution of this application. DETAILED DESCRIPTION

[0043] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way. The execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.

[0044] The HiSilicon Hi3516 platform has the characteristics of high computing performance, low power consumption, and small size, which facilitates the widespread deployment of cross-modal pedestrian re-identification algorithms in practice. Based on the HiSilicon Hi3516 platform, combined with the cross-modal pedestrian re-identification method, daytime and nighttime joint pedestrian re-identification is realized to create a 24-hour monitoring assistance system. The cross-modal pedestrian re-identification method based on the HiSilicon Hi3516 platform provides specific target pedestrian images, which can be used to query the monitoring trajectory of the target pedestrian during the day (or night) at night (or during the day). For example, for the target person, the RGB image is captured by the monitoring, and the target person hides until night. Using the cross-modal pedestrian re-identification method based on the HiSilicon Hi3516 platform, the daytime RGB image can be used to retrieve the infrared image taken by nighttime monitoring to obtain the nighttime monitoring trajectory of the target person.

[0045] The technical solution of this application utilizes the powerful image processing capability and high computing performance of HiSilicon Hi3516DV300, combined with the cross-modal pedestrian re-identification algorithm, to retrieve the pedestrian information of the video image in real time and extract the feature value, so as to realize offline or online cross-modal pedestrian re-identification; Figure 1 As shown, the embodiment method mainly includes steps S100-S400:

[0046] S100, obtaining an initial person re-identification model by training with a first data set;

[0047] Among them, the first data set is the data set trained before the pedestrian re-identification model is transplanted. In the embodiment, this data set can be constructed according to the SYSU-MM01 data and the RegDB data; the initial pedestrian re-identification model is the pedestrian re-identification model before the product platform is transplanted.

[0048] Specifically in the embodiment, firstly, the cross-modal person re-identification model (before transplantation) needs to be pre-trained, and the model is pre-trained in SYSU-MM01 and RegDB data respectively, such as Figure 2As shown, in the embodiment, the images in the training data set can be input into the model in two ways, one way is RGB image, and the other way is infrared image. The modality-specific features are first extracted, and then the RGB modality-specific features and the infrared modality-specific features are spliced ​​and input into the network to extract the modality common features. Then, pooling is also called downsampling processing. In the embodiment, two pooling methods can be used. Average pooling averages the feature points in a small neighborhood of the feature map, and the average value represents the small neighborhood; maximum pooling finds the maximum value of the feature points in a small neighborhood of the feature map, and the maximum value represents the small neighborhood. The pooling operation can reduce the feature dimension while retaining the effective information of the feature map, that is, reduce the number of parameters. The modality common features can reduce the feature dimension after the maximum pooling or average pooling operation, and extract identity-related features and identity-independent features through two fully connected layers. The batch normalization layer and the fully connected layer are followed by the identity-related feature fully connected layer to strengthen the identity-related information, so that the pre-trained cross-modal pedestrian re-identification model has the ability to extract cross-modal pedestrian features.

[0049] S200, converting the initial pedestrian re-identification model to obtain a cross-modal pedestrian re-identification model, and transplanting the cross-modal pedestrian re-identification model to a target platform;

[0050] Specifically in the embodiment, the target platform is the HiSilicon Hi3516 platform. In the embodiment, the Caffe model can be converted into a wk file through the official platform tool, and the file can be inferred in the Hi3516DV300. The Caffe model is an intermediate form of the initial pedestrian re-identification model converted to a cross-modal pedestrian re-identification model.

[0051] Regarding the process of transplanting the initial pedestrian re-identification model to the HiSilicon Hi3516 platform in the solution, in some feasible embodiments, the initial pedestrian re-identification model is converted to obtain a cross-modal pedestrian re-identification model, and the step S200 of transplanting the cross-modal pedestrian re-identification model to the target platform may include steps S210 and S220:

[0052] S210, performing model framework verification on the initial person re-identification model;

[0053] S220, determining that the initial person re-identification model does not belong to the Caffe framework, and converting the initial person re-identification model into the cross-modal person re-identification model under the Caffe framework;

[0054] Specifically in the embodiment, since the target platform, that is, the HiSilicon Hi3516 platform, only supports Caffe model transplantation, when the cross-modal pedestrian re-identification model is not a Caffe framework, the OONX format is used as an intermediate conversion layer to convert the cross-modal pedestrian re-identification model into the ONNX open source format, and then converted to a Caffe model. Then, the Caffe model is converted into a wk file using the official RuyStudio tool, which can be used for reasoning in Hi3516DV300.

[0055] S300, collecting real-time images through the target platform, and extracting pedestrian detection results from the real-time images;

[0056] Specifically in the embodiment, the preset yolov3 of HiSilicon Hi3516 can be called for real-time pedestrian detection, the sensor of HiSilicon Hi3516DV300 platform can be used for real-time image acquisition, the image is sent to the NNIE deep learning convolutional neural network hardware acceleration unit, and the HI_MPI_SVP_NNIE_LoadModel of the NNIE interface is used to load the preset yolov3 model of the platform. In the real-time image acquisition process, the HI_MPI_SYS_MmzAlloc_Cached interface is first used to allocate memory for storing yolov3 quantization model files in advance, and the HI_MPI_SVP_NNIE_LoadModel loads the yolov3 model from the memory allocated in advance by the user. The HI_MPI_SVP_NNIE_GetTskBufSize interface is used to parse the model segmentation, input nodes, output nodes, and running memory that needs to be allocated. Then, the running memory is allocated through HI_MPI_SYS_MmzAlloc_Cached, and the actual model inference calculation is performed through the interface HI_MPI_SVP_NNIE_Forward. In the calculation process, the neural network hardware acceleration of HiSilicon Hi3516DV300 is used, and HI_MPI_SVP_NNIE_ForwardWithBbox is used to score the pedestrian target prediction bounding box obtained by the yolov3 model and readjust the position to obtain the position information that meets the characteristics of pedestrians in the image.

[0057] S400, inputting the pedestrian detection result and the preset target into the cross-modal pedestrian re-identification model to obtain a pedestrian re-identification result;

[0058] In an embodiment, feature extraction is performed through a cross-modal pedestrian re-identification model and compared with the preset target features in the real-time collected image in step S300, and the recognition result is output. More specifically, the embodiment can apply the NNIE interface HI_MPI_SVP_NNIE_LoadModel to load the transplanted cross-modal pedestrian re-identification pre-trained model, and extract the cross-modal pedestrian re-identification features of the preset target query object through the interfaces HI_MPI_SVP_NNIE_Forward and HI_MPI_SVP_NNIE_ForwardWithBbox. The cross-modal pedestrian re-identification features of the pedestrian image detected by yolov3 are extracted through the interfaces HI_MPI_SVP_NNIE_Forward and HI_MPI_SVP_NNIE_ForwardWithBbox. The cross-modal features of the pedestrian image detected by yolov3 are compared with the cross-modal pedestrian re-identification features of the preset target query object, and the similarity result is output.

[0059] In some feasible embodiments, in step S100 of training the first data set to obtain an initial person re-identification model, the method may include steps S110-S140:

[0060] S110, acquiring an RGB image and an infrared image from the first data set;

[0061] S120, performing feature extraction on the RGB image to obtain RGB modal features, and performing feature extraction on the infrared image to obtain infrared modal features;

[0062] S130, combining the RGB modal feature and the infrared modal feature to obtain a mixed modal feature;

[0063] S140. Obtain a cross-modal feature according to the mixed modal feature output.

[0064] Specifically in the embodiment, the target cross-modal pedestrian re-identification network model includes a ResNet50-like backbone network, in which the first convolution module of the original ResNet50 network is separated and becomes two branches to extract modality-specific features for RGB images and infrared images, namely, the corresponding RGB modality features and infrared modality features, and then the two features are batch spliced. When the feature dimension of the batch RGB image is (b**h*w), the feature dimension of the batch infrared image is (b**h*w), where b is the number of images input to the network in this batch, c is the number of feature channels, and h and w are feature dimensions, the feature dimension after batch splicing of the two features is (2b*c*h*w), and the features after batch splicing include both RGB modality features and infrared modality features. The batch spliced ​​features are input into the remaining 4 convolution modules of the ResNet50 network. The ResNet50 network can be guided by artificial labels during training to overcome the differences in images and extract robust common pedestrian features between different images. Since the input of the network is mixed modal features, during the training process the network will learn to overcome the differences in image modalities and extract the common features between the two modal images to complete the task, so as to achieve the effect of extracting cross-modal features.

[0065] Furthermore, in an embodiment, the step S140 of obtaining a cross-modal feature according to the mixed modal feature output is specifically as follows: obtaining a cross-modal feature through a pooling layer output according to the mixed modal feature; wherein the cross-modal feature includes: identity-related features and identity-independent features; specifically in an embodiment, through a pooling layer in the model, and two fully connected layers that respectively extract identity-related features and identity-independent features, followed by a batch normalization layer and a fully connected layer, finally outputting the cross-modal feature.

[0066] In some feasible embodiments, in the process of obtaining identity-related features and identity-independent features through the pooling layer output according to the mixed modal features, the embodiment may include steps S131 and S132:

[0067] S131, obtaining a first feature distance of two images that do not belong to the same category, and a second feature distance of two images that belong to the same category in the mixed modal features;

[0068] S132, determining the identity-independent feature by using an inverse triplet loss function according to the first feature distance and the second feature distance;

[0069] Specifically in the embodiment, identity-related features can be extracted through two fully connected layers, including but not limited to correlation coefficient coordination and center clustering loss functions and triplet loss function guidance; while identity-independent features can be guided by the reverse triplet loss function, and the identity-related information can be strengthened by a batch normalization layer and a fully connected layer immediately after the identity-related feature fully connected layer, which can be guided by the cross-entropy loss function.

[0070] More specifically, the embodiment uses the reverse triplet loss to extract identity-independent features in the cross-modal features. The reverse triplet loss function is as follows:

[0071] Loss Tri-reverse =max(d a,n -d a,p +α,0)

[0072] Among them, Loss Tri-reverse is the reverse triplet loss value, d a,n is the first feature distance between two images that do not belong to the same class, d a,p is the feature distance between two images belonging to the same category, and α is the reverse triplet loss value Loss Tri-reverse The reverse triplet loss function is used to guide the network to extract identity-independent features, so that the feature distance between two images whose features do not belong to the same category is smaller than the feature distance between two images whose features belong to the same category, and cannot be used to judge identity.

[0073] In some feasible embodiments, in the process of obtaining identity-related features and identity-independent features through the pooling layer output according to the mixed modal features, the embodiment may further include step S133:

[0074] S133, extracting from the mixed modal features through correlation coefficient synergy and center clustering loss function, triple loss function and cross entropy loss function;

[0075] Specifically in the embodiment, the correlation coefficient synergy and the center clustering loss function are combined with the triple loss function and the cross entropy loss function to extract identity-related features in the cross-modal features. Exemplarily, the correlation coefficient synergy and the center clustering loss function are as follows:

[0076]

[0077] Among them, λ 1 and λ 2 is a hyperparameter, c is the number of classes in a mini-batch, is the RGB feature center vector of the i-th image in the small batch, is the infrared feature center vector of the i-th image in the small batch; is the relationship coefficient matrix of the i-th RGB feature in a small batch, is the relationship coefficient matrix of the i-th type of infrared features in a small batch, i = 1, 2, 3, …, n, and n is a positive integer.

[0078] Furthermore, after step S140 of obtaining the cross-modal feature through the output of the pooling layer according to the mixed modal feature, the method may further include steps S150-S160:

[0079] S150, performing orthogonal decoupling on the first eigenvector of the identity-related feature and the second eigenvector of the identity-independent feature through an orthogonal loss function;

[0080] S160, separating identity-irrelevant information from the identity-related features according to the result of orthogonal decoupling;

[0081] Specifically in the embodiment, an orthogonal loss function is used to orthogonally decouple identity-related features from identity-irrelevant features. The concept of orthogonality comes from vector orthogonality. When the inner product of two vectors is 0 or the angle between the two vectors is 90 degrees, the two vectors are called orthogonal. Mathematically, two vectors are considered to be orthogonal, that is, the two vectors are independent of each other. When the identity-related features and the identity-irrelevant features are regarded as two vectors, they are orthogonally decoupled so that the identity-related features and the identity-irrelevant features are independent of each other, that is, the two features are projected into mutually perpendicular feature spaces, which can separate and decouple the identity-irrelevant information in the identity-related features, and further improve the effectiveness of the identity-related features. More specifically, the orthogonal loss function in the embodiment is:

[0082]

[0083] Where b is the number of mini-batch inputs, is the identity-independent feature of the i-th image in the mini-batch input and The identity-related features of the ith image in the mini-batch input. The embodiment adopts a vector orthogonal method for calculation. When the loss function approaches 0 during network training, the two vectors are orthogonal, that is, the identity-independent features and the identity-related features are orthogonal.

[0084] Combined with Figure 3 , a complete process description of the technical solution of this application is as follows:

[0085] First, the cross-modal pedestrian re-identification model is pre-trained. The model is pre-trained in SYSU-MM01 and RegDB data respectively. Two images are input respectively, one is RGB image and the other is infrared image. The modality-specific features are extracted first, and then the RGB modality-specific features and infrared modality-specific features are spliced ​​and input into the network to extract the common features of the modalities. Pooling is also called downsampling. There are two main types of pooling. Average pooling averages the feature points in a small neighborhood of the feature map, and the average value represents the small neighborhood; maximum pooling finds the maximum value of the feature points in a small neighborhood of the feature map, and the maximum value represents the small neighborhood. The pooling operation can reduce the feature dimension while retaining the effective information of the feature map, that is, reduce the number of parameters. The common features of the modalities can reduce the feature dimension after the maximum pooling or average pooling operation, and extract identity-related features (guided by the correlation coefficient synergy and center clustering loss function, and triplet loss function) and identity-irrelevant features (guided by the reverse triplet loss function) through two fully connected layers. The batch normalization layer and the fully connected layer are followed by the identity-related feature fully connected layer to strengthen the identity-related information (guided by the cross entropy loss function). The pre-trained cross-modal pedestrian re-identification model has the ability to extract cross-modal pedestrian features.

[0086] Then, the cross-modal pedestrian re-identification model is converted to Caffe: HiSilicon Hi3516 platform only supports Caffe model transplantation. When the cross-modal pedestrian re-identification model is not a Caffe framework, the OONX format is used as the intermediate conversion layer to convert the cross-modal pedestrian re-identification model to the ONNX open source format, and then converted to a Caffe model.

[0087] Furthermore, the Caffe model was ported to the HiSilicon Hi3516 platform: the Caffe model was converted into a wk file using the official RuyStudio tool, which can be used for inference in the Hi3516DV300.

[0088] Furthermore, the preset yolov3 of HiSilicon Hi3516 is called for pedestrian detection: the sensor of HiSilicon Hi3516DV300 platform is used for real-time image acquisition, and the image is sent to the NNIE deep learning convolutional neural network hardware acceleration unit, and the HI_MPI_SVP_NNIE_LoadModel of the NNIE interface is used to load the preset yolov3 model of the platform. First, use the HI_MPI_SYS_MmzAlloc_Cached interface to allocate memory for storing yolov3 quantization model files in advance, HI_MPI_SVP_NNIE_LoadModel loads the yolov3 model from the memory allocated in advance by the user, and use the HI_MPI_SVP_NNIE_GetTskBufSize interface to parse the model segmentation, input nodes, output nodes, and running memory that needs to be allocated. Then use HI_MPI_SYS_MmzAlloc_Cached to allocate running memory, and use the interface HI_MPI_SVP_NNIE_Forward to perform actual model inference calculations. In the calculation process, the neural network hardware acceleration of HiSilicon Hi3516DV300 is used, and HI_MPI_SVP_NNIE_ForwardWithBbox is used to score the pedestrian target prediction bounding box obtained by the yolov3 model and readjust the position to obtain the position information that meets the characteristics of pedestrians in the image.

[0089] Finally, the cross-modal pedestrian re-identification model extracts features and compares them with the preset target features, and outputs the recognition results: Use the NNIE interface HI_MPI_SVP_NNIE_LoadModel to load the transplanted cross-modal pedestrian re-identification pre-trained model, and extract the cross-modal pedestrian re-identification features of the preset target query object through the interfaces HI_MPI_SVP_NNIE_Forward and HI_MPI_SVP_NNIE_ForwardWithBbox. Extract the cross-modal pedestrian re-identification features of the pedestrian image detected by yolov3 through the interfaces HI_MPI_SVP_NNIE_Forward and HI_MPI_SVP_NNIE_ForwardWithBbox. Compare the cross-modal features of the pedestrian image detected by yolov3 with the cross-modal pedestrian re-identification features of the preset target query object, and output the similarity result.

[0090] From the above specific implementation process, it can be concluded that the technical solution provided by the present invention has the following advantages or strengths compared with the prior art:

[0091] The technical solution of the present application provides a cross-modal pedestrian re-identification method based on the HiSilicon Hi3516 platform. The platform model used in the present invention is the HiSilicon Hi3516DV300. The solution utilizes the powerful image processing capability and high computing performance of the HiSilicon Hi3516DV300, combined with the cross-modal pedestrian re-identification algorithm, to retrieve the pedestrian information of the video image in real time and extract the feature values, thereby realizing offline or online cross-modal pedestrian re-identification, reducing the CPU occupancy rate, speeding up the cross-modal pedestrian re-identification speed, and maintaining the characteristics of low power consumption and small size.

[0092] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.

[0093] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0094] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor or other system that can fetch instructions from an instruction execution system, apparatus or device and execute instructions), or used in combination with these instruction execution systems, apparatuses or devices.

[0095] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0096] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

[0097] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A cross-modal person re-identification method, characterized in that: The following steps are involved: An initial person re-identification model is obtained by training the first data set; Converting the initial person re-identification model to obtain a cross-modal person re-identification model, and transplanting the cross-modal person re-identification model to a target platform; Real-time image acquisition is performed through the target platform, and pedestrian detection results are extracted from the real-time image; Inputting the pedestrian detection result and the preset target into the cross-modal pedestrian re-identification model to obtain a pedestrian re-identification result; The step of obtaining an initial person re-identification model by training the first data set includes: Acquire an RGB image and an infrared image from the first data set; Performing feature extraction on the RGB image to obtain RGB modal features, and performing feature extraction on the infrared image to obtain infrared modal features; The RGB modal feature and the infrared modal feature are combined to obtain a mixed modal feature; Obtaining a cross-modal feature according to the mixed modal feature output; The step of obtaining a cross-modal feature according to the mixed modal feature output includes: Obtaining cross-modal features through a pooling layer output according to the mixed modal features; the cross-modal features include: identity-related features and identity-independent features; The step of obtaining identity-related features and identity-independent features through a pooling layer output according to the mixed modal features includes: Obtaining a first feature distance of two images that do not belong to the same class and a second feature distance of two images that belong to the same class in the mixed modal feature; Determining the identity-independent feature by using an inverse triplet loss function according to the first feature distance and the second feature distance; The step of obtaining identity-related features and identity-independent features through a pooling layer output according to the mixed modal features includes: Extracted from the mixed modal features through correlation coefficient synergy and center clustering loss function, triple loss function and cross entropy loss function; After the step of obtaining the cross-modal feature through the pooling layer output according to the mixed modal feature, the method further includes the following steps: orthogonally decoupling the first eigenvector of the identity-related feature and the second eigenvector of the identity-independent feature through an orthogonal loss function; Separating identity-irrelevant information from the identity-relevant features according to the result of orthogonal decoupling; The step of converting the initial person re-identification model to obtain a cross-modal person re-identification model and transplanting the cross-modal person re-identification model to a target platform includes: Performing a model framework verification on the initial person re-identification model; It is determined that the initial person re-identification model does not belong to the Caffe framework, and the initial person re-identification model is converted into the cross-modal person re-identification model under the Caffe framework.

2. The cross-modal person re-identification method according to claim 1, characterized in that: The reverse triplet loss function is: Loss Tri-reverse =max(d a,n -d a,p +α,0); Among them, Loss Tri-reverse is the reverse triplet loss value, d a,n is the first feature distance between two images that do not belong to the same class, d a,p is the feature distance between two images belonging to the same category, and α is the reverse triplet loss value Loss Tri-reverse The boundary constant.

3. The cross-modal person re-identification method according to claim 1, characterized in that: The correlation coefficient synergy and center clustering loss function is: Among them, λ1 and λ2 are hyperparameters, c is the number of classes in a mini-batch, is the RGB feature center vector of the i-th image in the small batch, is the infrared feature center vector of the i-th image in the small batch; is the relationship coefficient matrix of the i-th RGB feature in a small batch, is the relationship coefficient matrix of the i-th type of infrared features in a small batch, i=1,2,3,…,n, where n is a positive integer.

4. The cross-modal person re-identification method according to claim 1, characterized in that: The orthogonal loss function is: Where b is the number of mini-batch inputs, is the identity-independent feature of the i-th image in the mini-batch input and Identity-related features for the i-th image in the input mini-batch.

Citation Information

Patent Citations

  • Cross-modal pedestrian re-identification method based on adaptive pedestrian alignment

    CN112651262A