A Pedestrian Re-identification Method Based on the Fusion of Local and Global Features with Mapping Relationships

By integrating local and global features in pedestrian recognition and optimizing weight allocation, the problem of insufficient accuracy of pedestrian recognition in cross-time and space scenarios is solved, and higher matching accuracy and stability are achieved.

CN116524566BActive Publication Date: 2025-07-25XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310496818.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-05
Publication Date
2025-07-25
Estimated Expiration
2043-05-05

AI Technical Summary

Technical Problem

In the pedestrian recognition in cross-time and space scenarios, traditional algorithms are difficult to effectively utilize local and global features, resulting in insufficient matching accuracy, especially in different environments, shooting angles and shooting distances.

Method used

The local and global feature fusion method based on mapping relationships is adopted, and the video is locally detected and segmented through the CGL-Net model, the local feature weight is calculated and summed with the global feature. The weight allocation is optimized using linear, segmented and nonlinear function mapping relationships, and finally the model is optimized through cross entropy and difficult ternary loss.

Benefits of technology

The accuracy of pedestrian re-identification in cross-time and space scenarios has been improved, and the matching performance under different conditions has been improved, which has performed better than existing methods, especially on the AST dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524566B_ABST
    Figure CN116524566B_ABST
Patent Text Reader

Abstract

The present invention discloses a pedestrian re-identification method based on the fusion of local and global features with a mapping relationship, belonging to video pedestrian re-identification. The present invention proposes a pedestrian re-identification method based on the fusion of local and global features with a mapping relationship, which is applied to cross-time and space scenarios to improve the accuracy of person matching. The present invention proposes a feature fusion network that fuses local image features and global image features by weights, and at the same time uses a conversion strategy for the proportion of local images to local feature weights under different function mapping relationships. Verified by the current cross-time and space cross-dressing pedestrian re-identification dataset, the performance of the present invention is superior to the best methods in the current video pedestrian re-identification field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of video pedestrian re-identification, and in particular, to a pedestrian re-identification method based on the fusion of local and global features with a mapping relationship. Background Art

[0002] Due to the factors of medium and long distances, the distribution of pedestrian bodies on image and video frame trajectories is not uniform and regular. Therefore, it is unreasonable to solely pursue the global features of the overall image by improving the feature extraction accuracy of convolutional operations on the overall image or by auxiliary methods such as multi-granularity to compare similarities, etc.

[0003] In addition, due to the reasons of across time and space, there are relatively complex changes in the appearance and behavior characteristics of the target person, such as the different spatial distributions of the human torso caused by behavioral actions, and the limitations of pure image feature extraction caused by irresistible factors such as clothing color, body posture, and age. Therefore, finding the key factors that can identify the person ID is the most crucial. Under this background condition, the difficulty of pedestrian re-identification has increased a lot. From the perspective of video-based pedestrian re-identification, the application of traditional excellent algorithms based on cross-dressing across time and space datasets is not ideal. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a pedestrian re-identification method based on the fusion of local and global features with a mapping relationship.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] A pedestrian re-identification method based on the fusion of local and global features with a mapping relationship, comprising the following steps:

[0007] (1) Establish a local and global feature fusion model CGL-Net based on mapping relationship transformation;

[0008] (2) Input the video into CGL-Net, perform local detection and segmentation on each frame of the video to obtain the local image corresponding to the head of the target person, and then use TCLNet to extract video features from the local part of the head in the original video sequence and the original video sequence respectively;

[0009] (3) Establish a mapping relationship between the local proportion of the head part of the target person in the original video and the weight distribution of the local feature tensor of the video, and obtain the corresponding weight based on the proportion of the local image of the person's head in the tracklet and the mapping relationship;

[0010] (4) Fuse the local video features of the person's head in the original image and the overall video features of the original video in an add manner according to the weights to obtain new outputs and features. The new outputs and features are successively calculated by the cross-entropy loss and the hard triplet loss to jointly calculate the final loss to optimize the parameters of the CGL-Net model during training.

[0011] Further, in step (1), the CGL-Net is based on the feature fusion of the baseline of the video-based Re-ID model method.

[0012] Further, in step (2), the open-pose pose point detection method is used to detect the head part structure and process it to segment the corresponding head map.

[0013] Further, in step (3), the established mapping relationship is a linear mapping:

[0014] f(x) ∝ x ==> f(x) = x, x ∈ (0, 1).

[0015] Where x is the local proportion of the target person's head part in the original video;

[0016] f(x) is the weight of the local features under the video;

[0017] 1 - f(x) is the weight of the global features under the video.

[0018] Further, in step (3), the established mapping relationship is a piecewise function:

[0019]

[0020] Where x is the local proportion of the target person's head part in the original video;

[0021] f(x) is the weight of the local features under the video;

[0022] 1 - f(x) is the weight of the global features under the video;

[0023] v is the magnification factor.

[0024] Further, in step (3), the established mapping relationship is a non-linear relationship:

[0025]

[0026]

[0027] Where x is the local proportion of the target person's head part in the original video;

[0028] f(x) is the weight of the local features of the video;

[0029] 1 - f(x) is the weight of the global features of the video;

[0030] v is the magnification factor.

[0031] Furthermore, in step (3), the established mapping relationship is a non - linear function transformation based on the magnification factor:

[0032] First, it is magnified in a direct - proportion relationship by the magnification factor of 20 according to the proportion of the local image, then:

[0033] W local = v * x, W global = 1 - x, x ∈ (0, 1), v = 20.0

[0034] Among them, W local represents the local feature weight, and W glocal represents the global feature weight;

[0035] The re - distributed weight is calculated through the calculation method of re - distributing the weight. Finally:

[0036]

[0037]

[0038] Furthermore, in step (3), the calculation process of the local proportion of the head part of the target person in the original video is as follows:

[0039] For the i - th image frame in the video trajectory, the proportion P i of the specific head part area of the target person in the frame is:

[0040]

[0041] Among them, Y max and Y min represent the upper and lower boundary coordinates of the head image in the original image, X max and X min represent the left and right boundary coordinates of the head image in the original image, and S original represents the size of the original image frame;

[0042] The proportion of the local part of the person's head in the video trajectory is represented by adding the combined head proportion values and taking the average:

[0043]

[0044] Among them, It represents the average value obtained by adding the ratios of the head images corresponding to the original image frames of the target person in the video track, and N represents the total count of video frames in the video track.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] The present invention proposes a pedestrian re-identification method based on the fusion of local and global features with mapping relationships, which is applied to cross-time and space scenarios to improve the accuracy of person matching. The present invention proposes a feature fusion network that fuses local image features and global image features by weights, and at the same time uses a conversion strategy of mapping the proportion of local images to local feature weights under different function mapping relationships. Verified by the current cross-time and space cross-dressing pedestrian re-identification dataset, the performance of the present invention is superior to the best method in the current field of video pedestrian re-identification. The present invention mainly solves the problem that general algorithms cannot be adapted due to the irregular limb distribution that may exist in the person images in the video under different environments, shooting angles, shooting distances, etc. Due to these special situations, the video images of pedestrians do not necessarily present a complete torso and regular posture, which is obviously unreasonable for general algorithms that focus on the overall image features. The present invention adds local features to intervene and assist, which is of great help to the feature comparison of the overall pedestrian image and has greatly improved performance compared with the existing algorithms.

[0047] Further, the determination method of the four-point azimuth coordinates of the head local image of the target person in the sampled video frame is used to calculate the area ratio. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is the overall network structure diagram of the present invention;

[0049] Figure 2 is a schematic diagram of identifying and segmenting the head image of a person in an image frame;

[0050] Figure 3 is a schematic diagram of the local ratio calculation module in the present invention;

[0051] Figure 4 is a schematic diagram of the conversion of the local ratio to the local feature weight calculation in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0053] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0054] CGL-Net (Combine Global and Local Feature Fusion) is designed to combine global features and local features and apply them to video trajectories in video-based pedestrian re-identification, which improves the overall performance of the algorithm to a certain extent.

[0055] The following further describes the present invention in detail with reference to the drawings:

[0056] The CGL-Net model is a local and global feature fusion model based on mapping relationship transformation. CGL-Net (Combine Global and Local Feature Net) is a feature fusion method based on the current relatively excellent baseline of video-based pedestrian re-identification model methods. Its purpose is to perform specific prior processing on the original image frames in the AST dataset to obtain the partial image of the target person's head corresponding to the image frames. Subsequently, the local features of the head in the extracted original image frames are added to the overall features of the original image frames themselves in a certain way to obtain new video features to improve the performance indicators of pedestrian re-identification under different conditions. The overall architecture of CGL-Net is as Figure 1 shown.

[0057] Local part detection and segmentation are performed to obtain the head part of the target person, specifically as follows:

[0058] For the image frame data contained in each video track in the AST dataset, it is possible to obtain an image of the human head part in a certain way. The method adopted in the present invention is to use the openpose pose point detection method to detect the head part structure and perform processing to segment the corresponding head image. The specific head structure recognition application is implemented based on openpose. Moreover, some modifications and functional expansions are made on the basis of this algorithm, and the complete segmentation of the head part can be fully carried out for local feature extraction. In addition, when the image of the human head is successfully detected and segmented, the corresponding four-corner azimuth coordinates of the head of the target person in each corresponding original image frame can be generated simultaneously for subsequent head ratio calculation. The schematic diagram of the improved algorithm for recognizing and segmenting the head is as follows Figure 2 as shown.

[0059] As Figure 2 described by the algorithm process shown, for an image frame, the human backbone structure corresponding to the original image frame is detected through openpose. In the present invention, the eye part in the human head range is used as the reference point α, and the detection point of the shoulder and neck part is used as the reference point β. Denote the vertical coordinate value of α as α′, and similarly, the vertical coordinate value of β as β′. The difference between their vertical coordinates is H, that is:

[0060] |α' - β'| = H

[0061] Taking H as half of the side length of the intercepted square area (the length from the eye part to the shoulder and neck part is generally greater than or equal to half of the size length of the overall head part, and it is reasonable to use this as the reference distance), then using this length and the reference line as the reference area and reference orientation for intercepting the square, the intercepted area of the human head part corresponding to the original image frame is also determined accordingly. That is, taking the eye orientation point α as the reference reference point and extending H distances to the left, right, up, and down, the intercepted area formed is a square of (2×H)×(2×H). This can ensure that the head part can be completely intercepted without affecting the subsequent experimental process, and the possible error caused by the intercepted area possibly containing non-head parts has no impact; in addition, the four-corner coordinates of this head area can be conveniently obtained through the eye coordinates and the shape characteristics of the square. Denote the eye coordinates as (X, Y), then the four-corner coordinates A, B, C, D are (taking the lower left corner as the 0 point):

[0062] A: (X - H, Y + H) B: (X + H, Y + H)

[0063] C: (X - H, Y - H) D: (X + H, Y - H)

[0064] In summary, the present invention uses an algorithm improved based on OpenPose to identify and intercept the image of the target person's head part in the original image frame and its specific four-corner coordinate orientation in the original image. Then, for any image frame in a specific video trajectory, the proportion P of the area of the specific head part of the target person in this frame is as follows:

[0065]

[0066] Among them, Y max 、Y min represent the upper and lower boundary coordinates of the head image in the original image, X max 、X min represent the left and right boundary coordinates of the head image in the original image, S original represents the size of the original image frame, and the ratio calculation process is as Figure 3 shown.

[0067] For example, Figure 3 to calculate the local ratio of the head part of any image frame in the video trajectory, first, through head detection and image segmentation, the head image and specific coordinate values are obtained respectively, and then the ratio of the local part in a single image is calculated, where P is the formula for calculating the area ratio based on the upper and lower limits of the coordinates.

[0068] Feature fusion is carried out as follows:

[0069] In this data set, the behavior trajectories of people contain information about the whole torso or part of the upper torso. Then, leveraging their common identifiable person ID information is the key information of the head part;

[0070] Secondly, in addition to the key head information, the global information of the whole person image can also provide certain environmental information assistance, which requires convolution or a relatively excellent feature extraction model. Therefore, the preliminary experiment on mid- and long-distance person re-identification is based on the feature fusion of global and local features. The feature fusion in this experiment is mainly carried out by adding multiple feature tensors.

[0071] The feature fusion strategy of CGL-Net. The specific feature fusion method of CGL-Net is obtained by adding the local video features of the human head in the original image and the overall video features of the original video according to a certain weight ratio. Here, it is specified that the weights of the two feature parts should add up to 1.0. In order to consider the fairness and impartiality in the comparison of video-based pedestrian re-identification algorithm models, the present invention chooses to use the same video-based pedestrian re-identification algorithm model: TCL-Net to extract video features at the local and global levels. For the specific calculation method of weight allocation, the area ratio of the head of the target person in the obtained video in the original image is used. Specifically for a certain video trajectory, the ratio of the local part of the human head in the video trajectory is represented by adding and averaging the head ratio values in each selected image frame:

[0072]

[0073] Among them, among them represents the average value obtained by adding the ratios of the specific head images corresponding to the original image frames of the target person in the video trajectory, and N represents the total count of video frames in the video trajectory.

[0074] Such as Figure 4 As shown, this process is used to calculate the weight of the local part of a specific video trajectory. When the ratios of the head parts of the target person in all selected image frames in a video trajectory are obtained, the corresponding average ratio will also be obtained by adding and averaging. This average ratio will be used as the subsequent f(x) parameter for actual weight calculation. After obtaining the specific local head ratio value of the video trajectory, through a certain function mapping relationship, that is, f(x) transformation, the specific weight ratio ξ that the local part feature of the head of the target person in a specific video feature occupies in the overall feature addition process can be obtained:

[0075]

[0076] Therefore, for the final feature calculated for a certain video trajectory, its final feature is obtained by adding the local video feature F local and the global video feature F_global through a certain ratio allocation method:

[0077] F = ξ × F local + (1.0 - ξ) × F global (0 ≤ ξ ≤ 1.0).

[0078] The reasons for choosing to use weight addition at the video feature tensor level rather than weight addition for each image frame feature in the video are as follows:

[0079] 1) For any algorithm model, it is not always possible to determine whether each video frame of a video trajectory in a person re-identification dataset has been applied. Each model and algorithm differs in data selection and sample sampling strategies, including full sampling, equal-interval sampling, random sampling, etc. The randomness is relatively large. Therefore, it is impossible to determine at the image frame level which specific image frames a certain algorithm model has applied to in terms of sampling, and thus it is impossible to determine which specific image frames to use for the corresponding head proportion for subsequent mathematical calculations.

[0080] 2) In the process of constructing a person re-identification dataset, the specific information of a video trajectory is generally represented by extracting video frames from the video trajectory. Obviously, for videos of different time lengths, there are differences in the number of video frames in a set of video trajectory image data (even if a piecewise function type of case-by-case discussion is adopted, it is still impossible to ensure 100% consistency in quantity). For the same video trajectory, due to the principles of short-term and locality, the difference in the proportion of the target person's head image in the image frames is not significant. Therefore, it can be considered reasonable to use the average value of the proportion of the person's head in this video trajectory as the weight basis for this video feature.

[0081] 3) If it is necessary to consider the proportion situation specifically at the level of specific image frames in the video, it can also be achieved. However, due to different image data screening and sampling strategies and different selection and frame extraction methods for the number of video frames in a single video trajectory, the specific algorithm code implementation will also be continuously modified and changed accordingly. This is obviously very troublesome and not applicable. The algorithm has relatively high time complexity and space complexity. Considering the average value at the video trajectory level can undoubtedly be obtained in advance through data preprocessing, and then can be directly stored in the form of a record document for subsequent use, saving the waste of memory and video memory.

[0082] Based on the function mapping setting from local proportion to feature weight allocation, it is as follows:

[0083] For the mapping relationship from the local proportion of the target person's head part in the original video to the specific weight allocation of the feature tensor under this video, a specific mapping function needs to be designed to support it. Moreover, it is necessary to ensure that this function can actively improve the performance of the algorithm model on the AST dataset after feature fusion. Then this section will start from different types of function mappings and compare the relationship between local proportion and feature weight under different mapping relationships.

[0084] Linear mapping

[0085] If the weight ratio of the head part is used as a benchmark, it is obvious that as the proportion of the head part in the image increases, the weight of the feature vector of the head part in the overall feature fusion should be greater, because at this time, as the proportion of the head part increases, the influence of environmental factors brought by the so-called global image will gradually decrease (you can try to think about the influence brought by the head part gradually increasing theoretically from a whole torso to a headshot).

[0086] Based on this, obviously, linear mapping can meet the above conditions. For linear mapping, a relatively classic and direct mapping relationship is the proportional relationship mapping, which can well satisfy the prior assumption that "as the local proportion increases, the influence brought by this part should also gradually increase", that is:

[0087] f(x)∝x ==> f(x) = x, x∈(0,1).

[0088] The domain of the proportional mapping relationship is (0,1), because for the proportion of the local part, it is obviously a fraction or decimal between 0 and 1.0.

[0089] Analyzed from a mathematical sense, the proportional relationship is relatively simple in its own nature. f(x) increases as x increases, and it is differentiable and integrable within the domain (0,1). There are no subtle problems in actual code applications and experiments that cause errors. At the same time, it also well satisfies the "increasing relationship" from the proportion relationship to the feature weight mapping, and is relatively simple in calculation without complex space and time consumption.

[0090] Relevant code experiments were conducted to prove this relationship. The experimental results are shown in Table 1. From the experimental results, it can be seen that linear mapping (proportional) does improve the performance in terms of feature summation, which also proves that the mapping relationship direction from local proportion to feature weight is reasonable and feasible.

[0091] Table 1 Experimental results when using linear mapping

[0092]

[0093] Piecewise function

[0094] The linear relationship can provide certain support for performance improvement. However, for the local part of the target person's head, in fact, as long as the head part of the image can be intercepted, the information it can bring is also sufficient, and even the person ID can be directly confirmed. Then, for the head part with a relatively small proportion, in order to appropriately enhance the influence factor of the head part under a small proportion, an amplification parameter can be selected as an aid to increase the weight ratio under a small proportion. However, because the global and local weight ratios should add up to 1.0, in order to prevent the data of the part with a larger proportion from exceeding 1.0 after amplification, the upper limit needs to be set manually, that is, 1.0 is the upper limit of the local feature weight.

[0095] Based on this, an actual piecewise function can be designed to optimize the limitations brought by the linear function:

[0096]

[0097] Among them, v represents the specific addition parameter for the proportion, that is, the amplifier.

[0098] For the specific value of v, there is no specific reference value as a benchmark. Therefore, in this invention, through a large number of experimental verifications, the v corresponding to the relatively excellent performance is determined as the promotion effect of the piecewise function discussed in this section on the actual performance improvement of the mapping from "local proportion to local part feature weight" on the AST dataset.

[0099] Table 2 Experimental results when using the piecewise function

[0100]

[0101] As shown in Table 2, the data brought by a large number of v-value transformation experiments show that as the v value increases, the improvement of the overall algorithm model performance brought by the change of the weight of the local feature part in feature fusion by the piecewise function exists, and the overall performance shows a trend of increasing first and then decreasing in terms of mAP and rank-1, that is, there is an extreme point among them, making this gain extremely large. In this invention, the specific extreme point v = 20.0 obtained in the experiment is used as the actual application value. Through actual experimental data verification, it can be seen that the performance gain brought by the piecewise function increases mAP and rank-1 by 2.89% and 10.30% respectively compared with the linear transformation, showing an obvious improvement.

[0102] Nonlinear relationships sigmoid & tanh

[0103] Although the changes and optimizations in the linear relationship do improve the performance of the overall feature fusion, there is still a possibility of a bottleneck. Specifically, the growth rate (i.e., slope) of the mapping relationship from local proportion to feature weight is consistent within any tiny local domain of definition, without any change. However, it is obvious that the influence of local parts of the image on the overall feature tensor should also vary under different degrees. Therefore, different slopes should exist for different x values to distinguish this. So, through the transition of linear transformation, the scope of non-linear mapping should be considered.

[0104] To make the function involved in the mapping relationship smoother and differentiable within the specified domain of definition (0, 1), it is natural to think of two relatively excellent smooth curves, the sigmoid and tanh non-linear activation functions, to participate in the experiment. Their greatest characteristic is that within the domain of definition (0, 1), the value of f(x) can be well restricted between (0, 1), and unlike piecewise functions, there is no need for fixed values to limit the upper bound range.

[0105] In this application of non-linear activation functions, a magnification factor consistent with that of the piecewise function is added to enhance the influence of "local parts with small proportions", that is, the mapping relationship is as follows:

[0106]

[0107]

[0108] Among them, v is the magnification factor. Since the selection strategy involved in the magnification factor in the piecewise function is the same, in this experiment, it is also found that when v = 5.0, the influence effect on the two non-linear functions is the best. Therefore, the experimental results are compared with other mappings based on v = 5.0 as the benchmark.

[0109] Table 3 Experimental results of two non-linear relationships and other mapping relationships when the magnification factor is 5.0

[0110]

[0111] By comparing the data in Table 3, the mapping of the sigmoid function has a relatively more positive influence on the local feature weights in feature fusion, that is, the performance is better. The reason may be that for local images with small proportions, sigmoid can better activate and amplify these image information to improve the prediction performance of the overall model (because it can be seen from the function image that at the starting value of x = 0, sigmoid can give more and larger weights compared to tanh, which is beneficial for the local image to better distinguish the person ID itself).

[0112] Non-linear function transformation based on the magnification factor

[0113] The normal linear transformation, i.e., the proportional function mapping, is the simplest function mapping relationship that satisfies the increasing relationship. However, this will lead to the situation where the weight of the local image features is negligible under a very small proportion; for the situation where the weight proportion brought by the local image under a small proportion is small, the solution of the piecewise function is to enhance the weight proportion of the small-proportion image through a magnification factor. This is reasonable because the local part of the image is needed to affect the final ID classification task.

[0114] Then, considering the combination of these two aspects, first, based on the normal proportional relationship, the feature weight X of the local part of the image and the weight 1.0 - x of the global image features are obtained from the local image proportion x. Since the small-proportion situation is considered, the local part weight X can be magnified accordingly. Here, the value of the magnification factor v = 20.0, which performed well in the piecewise function experiment verification, is used as a reference. Then, there will be:

[0115] W local = v * x, W global = 1 - x, x ∈ (0, 1), v = 20.0

[0116] where W local and W global represent the local feature weight and the global feature weight.

[0117] In addition, to consider the actual weight upper limit situation, the weights of the two parts cannot exceed the upper limit of 1.0, and at the same time, the sum of the two is 1.0. Therefore, the reallocated weight can be calculated by reallocating the weight calculation method:

[0118]

[0119]

[0120] Within the domain (0, 1), the function curve has the characteristics of being smooth, differentiable, and integrable, similar to the non-linear activation function, meeting the requirements of feature fusion in the neural network model.

[0121] As shown in Table 4, through the comparison of the mAP and rank-1 metrics, it can be seen that the non-linear transformation designed inspired by the proportional transformation and the piecewise function has a significant improvement in mAP, with a 1.1% increase. However, rank-1 has a slight attenuation compared to the sigmoid function, indicating that this non-linear transformation has a positive effect on the improvement of the general performance of the overall data, and the first hit rate is not greatly affected, proving the feasibility of the non-linear mathematical method.

[0122] Table 4 Experimental results of non-linear function transformation based on magnification factor and other mappings

[0123]

[0124]

[0125] The above content is only for explaining the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.

Claims

1. A pedestrian re-identification method based on the fusion of local and global features with mapping relationship, characterized in that, It includes the following steps: (1) Establish a local and global feature fusion model CGL-Net based on mapping relationship transformation; (2) Input the video into CGL-Net, perform local detection and segmentation on each frame of the video to obtain the local image corresponding to the target person's head, and then use TCLNet to extract video features from the local part of the head in the original video sequence and the original video sequence respectively; (3) Establish a mapping relationship between the local proportion of the target person's head part in the original video and the weight allocation of the local feature tensor in the video, and obtain the corresponding weight based on the proportion of the local image of the person's head in the tracklet and the mapping relationship; (4) Fuse the local video features of the person's head in the original image and the overall video features of the original video in an "add" manner according to the weight to obtain new outputs and features, and the new outputs and features are used to jointly calculate the final loss through cross-entropy loss and hard triplet loss to optimize the parameters of the CGL-Net model during training.

2. The pedestrian re-identification method based on local and global feature fusion according to claim 1, wherein In step (1), the CGL-Net is based on the feature fusion of the Re-ID model method baseline for videos.

3. The pedestrian re-identification method based on local and global feature fusion according to claim 1, wherein In step (2), the open-pose pose point detection method is used to detect the head part structure and process it to segment the corresponding head map.

4. The pedestrian re-identification method based on local and global feature fusion according to claim 1, characterized in that In step (3), the established mapping relationship is a linear mapping: f(x) ∝ x ==> f(x) = x, x ∈ (0, 1). Where x is the local proportion of the target person's head part in the original video; f(x) is the weight of the local features in the video; 1 - f(x) is the weight of the global features in the video.

5. The person re-identification method based on local and global feature fusion according to claim 1, wherein In step (3), the established mapping relationship is a piecewise function: Where x is the local proportion of the target person's head part in the original video; f(x) is the weight of the local features in the video; 1 - f(x) is the weight of the global features in the video; v is the amplification factor.

6. The pedestrian re-identification method based on local and global feature fusion of mapping relationship according to claim 1, characterized in that In step (3), the established mapping relationship is a non-linear relationship: Where x is the local proportion of the target person's head part in the original video; f(x) is the weight of the local features in the video; 1 - f(x) is the weight of the global features in the video; v is the amplification factor.

7. The pedestrian re-identification method based on local and global feature fusion according to claim 1, characterized in that In step (3), the established mapping relationship is a non-linear function transformation based on the amplification factor: First, it is amplified by a factor of 20 in direct proportion to the local image proportion, then: W local = v * x, W global = 1 - x, x ∈ (0, 1), v = 20.0 Among them, W local represents the local feature weight, and W glocal represents the global feature weight; The reallocated weight is calculated through a weight reallocation calculation method, and finally:

8. The method for pedestrian re-identification based on local and global feature fusion with mapping relationship according to any one of claims 1-7, characterized in that In step (3), the calculation process of the local proportion of the target person's head part in the original video is: For the i-th image frame in the video trajectory, the proportion P of the area of the specific head part of the target person in the frame i is as follows: Among them, Y max and Y min represent the upper and lower boundary coordinates of the head image in the original image, X max and X min represent the left and right boundary coordinates of the head image in the original image, and S original represents the size of the original image frame; The proportion of the local part of the person's head in the video track is represented by taking the average of the added combined head proportion values: Among them, represents the average value obtained by adding the ratios of the head images corresponding to the original image frames of the target person in the video track, and N represents the total count of video frames in the video track.

Citation Information

Patent Citations

  • Video pedestrian re-identification method based on multi-scale feature fusion

    CN114299542A

  • Pedestrian re-identification method and system based on comparison features

    CN114429648A