Training Method for Pedestrian Matching Model and Cross-Resolution Pedestrian Matching Method
By improving the triple loss function and dynamically fusing high-frequency and low-frequency characteristics, the problem of low resolution matching accuracy in pedestrian matching is solved, and efficient and accurate pedestrian matching is achieved, which is suitable for real-time applications.
Patent Information
- Application Number
- CN202510280326.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing technology has the problem of low cross-resolution matching accuracy in the pedestrian matching process, and the existing methods are cumbersome and include many sub-networks, resulting in too long inference time and real-time cross-resolution matching of pedestrians cannot be achieved.
By calculating the high-frequency component energy and low-frequency component energy of the images in the multi-resolution image training dataset, the image resolution quality evaluation index is constructed, and the triple loss function is improved. The improved triple loss function is used to train the pedestrian matching model, and the high-frequency and low-frequency features are dynamically fused to improve matching accuracy.
It realizes that the matching accuracy and robustness of the pedestrian matching model can be improved without increasing the inference time, and can stably process images with different resolutions.
Smart Images

Figure CN119785163B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning, neural networks, and image processing, and particularly to a training method for a pedestrian matching model, a cross-resolution pedestrian matching method, an electronic device, and a storage medium. Background Art
[0002] Pedestrian matching is an emerging technology in the field of intelligent video analysis and belongs to the category of image processing and analysis in complex video environments. The main goal of this technology is to verify the identity of pedestrians in image sequences captured by different cameras and has a wide range of application scenarios. However, in the actual application of pedestrian matching, due to the inconsistent distances of the targets from the camera, the resolutions of the captured pictures are different. Low-resolution images contain much fewer identity details, and directly cross-resolving and matching image pairs will lead to a significant performance decline. Existing technologies only consider the extraction and matching of image features in pedestrian matching and do not consider the problem of low cross-resolution matching accuracy in actual applications; or although the problem of cross-resolution matching is considered, the steps are cumbersome and involve many sub-networks, such as image super-resolution networks, etc., and the inference time is long in actual applications, and real-time pedestrian cross-resolution matching cannot be achieved; or when using triplet loss to train the model, the impact on the Euclidean distance calculation and hard sample mining in the triplet loss caused by different-resolution samples is not considered. Summary of the Invention
[0003] In view of the above problems, the present invention provides a training method for a pedestrian matching model, a cross-resolution pedestrian matching method, an electronic device, and a storage medium, which are used to solve at least one of the above technical problems.
[0004] According to a first aspect of the present invention, there is provided a training method for a pedestrian matching model, including:
[0005] Calculating the high-frequency component energy and low-frequency component energy of images in a multi-resolution image training dataset by using a pedestrian matching model, and improving a triplet loss function through an image resolution quality evaluation index constructed by the high-frequency component energy and the low-frequency component energy to obtain an improved triplet loss function;
[0006] Using the pedestrian matching model to extract features from the high-frequency component energy and the low-frequency component energy respectively, and dynamically fusing the extracted high-frequency features and low-frequency features to obtain a two-stream fusion feature;
[0007] Processing the two-stream fusion feature and the true value label corresponding to the two-stream fusion feature by using the improved triplet loss function to obtain a triplet loss value, and updating the parameters of the pedestrian matching model by using the triplet loss value;
[0008] Iteratively perform feature extraction and dynamic fusion operations, triplet loss value calculation operations, and model parameter update operations until the preset training conditions are met, and obtain a trained pedestrian matching model.
[0009] According to an embodiment of the present invention, the above-mentioned calculation of the high-frequency component energy and low-frequency component energy of the images in the multi-resolution image training dataset by using the pedestrian matching model includes:
[0010] Based on the authorization of the experimental user, collect the images of the experimental user by setting the downsampling ratio of the image acquisition device to obtain a multi-resolution image training dataset;
[0011] Perform wavelet transform on the images in the multi-resolution image training dataset through the pedestrian matching model to obtain wavelet transform coefficients;
[0012] Define an energy function based on the wavelet transform coefficients, and use the energy function to calculate the high-frequency component energy of the high-resolution images and the low-frequency component energy of the low-resolution images in the multi-resolution image training dataset respectively.
[0013] According to an embodiment of the present invention, the above-mentioned improvement of the triplet loss function by using the image resolution quality evaluation index constructed by the high-frequency component energy and the low-frequency component energy to obtain the improved triplet loss function includes:
[0014] Calculate the total energy between the high-frequency component energy and the low-frequency component energy, and perform an operation on the high-frequency component energy and the total energy to obtain the image resolution quality evaluation index;
[0015] By using the image resolution quality evaluation index as the weight of the triplet loss function to perform weighted processing on the feature data to be processed by the triplet loss function, the improved triplet loss function is obtained.
[0016] According to an embodiment of the present invention, the above-mentioned improvement of the triplet loss function by using the image resolution quality evaluation index as the weight of the triplet loss function to perform weighted processing on the feature data to be processed by the triplet loss function to obtain the improved triplet loss function includes:
[0017] Calculate the Euclidean distance matrix between the feature data, and use the image resolution quality evaluation index to calculate the image resolution quality score ratio of the images in the multi-resolution image training dataset;
[0018] Preprocess the image resolution quality score ratio, and use the preprocessed image resolution quality score ratio to obtain the resolution weight;
[0019] Operate on the resolution weights and the Euclidean distance matrix to obtain the Euclidean weighted distance matrix, and use the weighted distance matrix to obtain the farthest positive sample and the nearest negative sample of the images in the multi-resolution image training dataset;
[0020] Improve the triplet loss function by using the farthest positive sample, the nearest negative sample, and the Euclidean weighted distance matrix to obtain the improved triplet loss function.
[0021] According to an embodiment of the present invention, the above-mentioned feature extraction of the high-frequency component energy and the low-frequency component energy by using the pedestrian matching model respectively includes:
[0022] Perform wavelet decomposition on the low-frequency component energy by using the low-frequency identity flow network of the pedestrian matching model to obtain the low-frequency component energy subbands;
[0023] Extract the global structure information related to identity from the low-frequency component energy subbands by using the low-frequency identity flow network to obtain the low-frequency features.
[0024] According to an embodiment of the present invention, the above-mentioned feature extraction of the high-frequency component energy and the low-frequency component energy by using the pedestrian matching model respectively further includes:
[0025] Perform wavelet decomposition on the high-frequency component energy by using the high-frequency detail flow network of the pedestrian matching model to obtain the high-frequency component energy subbands;
[0026] Extract the resolution-sensitive local detail features from the high-frequency component energy subbands by using the high-frequency detail flow network to obtain the high-frequency features, thereby completing the high-frequency feature extraction and the low-frequency feature extraction based on wavelet-spatial domain two-stream feature decoupling.
[0027] According to an embodiment of the present invention, the above-mentioned dynamic fusion of the extracted high-frequency features and low-frequency features to obtain the two-stream fusion features includes:
[0028] Dynamically generate the coefficients of the resolution discriminator in the pedestrian matching model based on the image resolution quality evaluation index;
[0029] Based on the learnable gating mechanism, dynamically fuse the high-frequency features and the low-frequency features by using the coefficients of the resolution discriminator to obtain the two-stream fusion features.
[0030] According to the second aspect of the present invention, a cross-resolution pedestrian matching method is provided, including:
[0031] Calculate the high- and low-frequency component energies of the query image set by using the trained pedestrian matching model, and perform high- and low-frequency feature extraction based on wavelet-spatial domain two-stream feature decoupling on the high- and low-frequency component energies of the query image set to obtain the high- and low-frequency features of the query image set, wherein the trained pedestrian matching model is trained based on the above-mentioned training method of the pedestrian matching model;
[0032] The trained pedestrian matching model is used to dynamically fuse the high - and low - frequency features of the query image set based on a learnable gating mechanism to obtain the two - stream fusion features of the query image set;
[0033] The trained pedestrian matching model calculates the high - and low - frequency component energies of the reference image set, and performs high - and low - frequency feature extraction on the high - and low - frequency component energies of the reference image set based on wavelet - spatial domain two - stream feature decoupling to obtain the high - and low - frequency features of the reference image set, where the reference image set and the query image set have different resolutions;
[0034] The trained pedestrian matching model is used to dynamically fuse the high - and low - frequency features of the reference image set based on a learnable gating mechanism to obtain the two - stream fusion features of the reference image set;
[0035] The trained pedestrian matching model calculates the similarity between the two - stream fusion features of the query image set and the two - stream fusion features of the reference image set, and matches the pedestrians in the query image set and the reference image set based on the similarity to obtain the pedestrian matching result.
[0036] The third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the above - mentioned one or more processors execute the above - mentioned one or more computer programs to implement the steps of the above - mentioned method.
[0037] The fourth aspect of the present invention further provides a computer - readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above - mentioned method are implemented.
[0038] The training method of the pedestrian matching model provided by the present invention constructs an image resolution quality evaluation index by calculating the high - and low - frequency component energies in the training dataset, improves the triplet loss function using the image resolution quality evaluation index, and trains the pedestrian matching model using the improved triplet loss function, fully considering the influence of high - and low - resolution differences on the pedestrian image feature representation, so that the trained pedestrian matching model has stable performance in processing images with different resolutions, improves the inference accuracy and robustness of the model while ensuring a high inference speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Through the following description of the embodiments of the present invention with reference to the drawings, the above - mentioned content and other objects, features and advantages of the present invention will become clearer. In the drawings:
[0040] Figure 1 is an application scenario diagram of the cross - resolution pedestrian matching method according to an embodiment of the present invention;
[0041] Figure 2 is a flowchart of a method for training a pedestrian matching model according to an embodiment of the present invention;
[0042] Figure 3 is a flowchart of a cross - resolution pedestrian matching method according to an embodiment of the present invention;
[0043] Figure 4 is a structural diagram of a cross - resolution pedestrian matching system according to an embodiment of the present invention;
[0044] Figure 5 is a block diagram of an electronic device suitable for implementing the method for training a pedestrian matching model and the cross - resolution pedestrian matching method according to an embodiment of the present invention. Detailed implementation manners
[0045] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a thorough understanding of the embodiments of the present invention. However, obviously, one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well - known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0046] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. as used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0047] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0048] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0049] The research object of pedestrian matching is the entire characteristics of a person, which does not rely on face recognition and includes clothing, body posture, hairstyle, gesture, etc. This technology focuses on the problem of pedestrian re-identification across cameras and can be used as an extension of face recognition technology and applied to more scenarios. When performing pedestrian matching, generally a target pedestrian picture is given, and it is searched in the saved picture database and the real-time video stream, and finally the pedestrian matching situation is analyzed.
[0050] Existing technical solutions for pedestrian matching have technical problems such as a cumbersome inference network structure, too long inference time, and inability to be applied to scenarios of multi-resolution images.
[0051] The present invention provides a training method for a pedestrian matching model and a cross-resolution pedestrian matching method, which can solve the problem of low cross-resolution matching degree in the process of pedestrian matching without increasing the inference time and without affecting the inference efficiency.
[0052] Figure 1 It is an application scenario diagram of the cross-resolution pedestrian matching method according to an embodiment of the present invention.
[0053] As Figure 1 shown, the application scenario 100 according to this embodiment may include the technical fields of deep learning, neural networks, and image processing. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0054] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0055] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0056] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0057] It should be noted that the cross-resolution pedestrian matching method provided by the embodiments of the present invention can generally be executed by server 105. Correspondingly, the cross-resolution pedestrian matching device provided by the embodiments of the present invention can generally be set in server 105. The cross-resolution pedestrian matching method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the cross-resolution pedestrian matching device provided by the embodiments of the present invention can also be set in a server or a server cluster different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0058] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0059] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 The following will be based on Figures 2 to 3 the described scenarios, and will describe in detail the training method of the pedestrian matching model and the cross-resolution pedestrian matching method of the disclosed embodiments through
[0060] It should be particularly noted that: in the embodiments of the present application, certain industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present application, but it does not mean that the applicant has already or necessarily used this solution.
[0061] In addition, in the technical solutions of the present invention, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure, and application, etc. of the relevant data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0062] In the scenario of making automated decisions using personal information, the method, device, and system provided by the embodiments of the present invention all provide corresponding operation entrances for users to choose to agree or reject the results of automated decisions; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision" here refers to the activity of automatically analyzing and evaluating an individual's behavior habits, hobbies, or economic, health, credit status, etc. through a computer program and making a decision. The expression "expert decision" here refers to the activity of making a decision by a person who specializes in a certain field, has specialized experience, knowledge, and skills, and has reached a certain professional level.
[0063] Figure 2 It is a flowchart of the training method of the pedestrian matching model according to the embodiments of the present invention.
[0064] As Figure 2 shown, the above-mentioned training method of the pedestrian matching model includes operations S210 to S240.
[0065] In operation S210, the high-frequency component energy and low-frequency component energy of the images in the multi-resolution image training dataset are calculated using the pedestrian matching model, and the triple loss function is improved through the image resolution quality evaluation index constructed by the high-frequency component energy and low-frequency component energy, resulting in an improved triple loss function.
[0066] The above-mentioned multi-resolution image training dataset includes high-resolution images and low-resolution images of experimental users, and the resolution of the above-mentioned images can be set according to actual needs or scenario requirements.
[0067] The above-mentioned pedestrian matching model is constructed based on a classification backbone network and a pedestrian re-identification classifier. The classification backbone network, optionally, such as resnet, vit, etc.; those skilled in the art can adopt other neural networks or models according to actual needs or application scenario requirements.
[0068] In the process of calculating the high- and low-frequency component energies of the images in the multi-resolution image training dataset in the above-mentioned operation S210, wavelet transform is utilized. Those skilled in the art can adopt other methods to obtain the high- and low-frequency component energies of the images according to actual needs or application scenario requirements.
[0069] In operation S220, the pedestrian matching model is used to extract features from the high-frequency component energy and low-frequency component energy respectively, and the extracted high-frequency features and low-frequency features are dynamically fused to obtain dual-stream fusion features.
[0070] In the above operation S220, in the process of obtaining high and low frequency features, the wavelet-spatial domain two-stream feature decoupling method is utilized, that is, the wavelet transform is used to obtain the energy of high and low frequency components, and then the above high and low frequency component energies are processed respectively to obtain high and low frequency features.
[0071] In operation S230, the improved triplet loss function is used to process the two-stream fusion feature and the ground truth label corresponding to the two-stream fusion feature to obtain the triplet loss value, and the triplet loss value is used to update the parameters of the pedestrian matching model.
[0072] In operation S240, the feature extraction and dynamic fusion operation, the triplet loss value calculation operation, and the model parameter update operation are iteratively performed until the preset training conditions are met, and the trained pedestrian matching model is obtained.
[0073] The training method of the pedestrian matching model provided by the present invention constructs an image resolution quality evaluation index by calculating the high and low frequency component energies in the training data set, improves the triplet loss function by using the image resolution quality evaluation index, and trains the pedestrian matching model by using the improved triplet loss function, fully considering the influence of high and low resolution differences on the pedestrian image feature representation, so that the trained pedestrian matching model has stable performance in processing images with different resolutions, improves the inference accuracy of the model and the robustness of the model while ensuring a high inference speed.
[0074] According to an embodiment of the present invention, the above-mentioned calculation of the high frequency component energy and the low frequency component energy of the images in the multi-resolution image training data set by using the pedestrian matching model includes: based on the authorization of the experimental user, the images of the experimental user are collected by setting the downsampling ratio of the image acquisition device to obtain a multi-resolution image training data set; the pedestrian matching model performs wavelet transform on the images in the multi-resolution image training data set to obtain wavelet transform coefficients; an energy function is defined based on the wavelet transform coefficients, and the energy function is used to calculate the high frequency component energy of the high resolution images and the low frequency component energy of the low resolution images in the multi-resolution image training data set respectively.
[0075] In the process of obtaining the image data related to the experimental user, the permission and authorization of the experimental user are obtained.
[0076] The following further elaborates in detail the process of calculating the high frequency component energy and the low frequency component energy of the images in the multi-resolution image training data set through specific embodiments.
[0077] Using Market1501 to make a multi-resolution image training data set: Take the high resolution images taken by a camera Perform downsampling to obtain low resolution images , and the downsampling ratio is The random magnification in [[]], and the resolution of the high-resolution images captured by the remaining cameras remains unchanged. At this time, the images in the training set include high-resolution images and low-resolution images ; then downsample all the high-resolution images in the query set (query image set, the same below), and the downsampling magnification is the random magnification in [[]]; the resolution of the images in the gallery set (reference image set, the same below) remains unchanged.
[0078] Calculate the energy of the high-frequency components, the energy of the low-frequency components, and the ratio of the high-frequency energy to the total energy of the pictures in the multi-resolution image training dataset using wavelet transform: :
[0079] Perform wavelet transform on the pictures in the multi-resolution image training dataset to obtain wavelet transform coefficients , as shown in Equation (1):
[0080] (1),
[0081] where is the wavelet decomposition function, is the wavelet type, is the decomposition level.
[0082] Define an energy function for the energy of the wavelet transform coefficients , as shown in Equation (2):
[0083] (2),
[0084] where is a sub-band in the wavelet transform coefficients, and are the number of rows and columns of the sub-band.
[0085] The energy of the high-frequency components is the sum of the energies of all high-frequency sub-bands, as shown in Equation (3):
[0086] (3),
[0087] where is the high-frequency sub-band at the th layer.
[0088] The energy of the low-frequency components is the sum of the energies of all high-frequency subbands, as shown in Equation (4):
[0089] (4),
[0090] wherein, is the lowest-frequency subband.
[0091] According to an embodiment of the present invention, the above image resolution quality evaluation index constructed by the high-frequency component energy and the low-frequency component energy is used to improve the triplet loss function. The improved triplet loss function includes: calculating the total energy between the high-frequency component energy and the low-frequency component energy, and performing an operation on the high-frequency component energy and the total energy to obtain the image resolution quality evaluation index; performing weighted processing on the feature data to be processed by the triplet loss function by using the image resolution quality evaluation index as the weight of the triplet loss function, so as to obtain the improved triplet loss function.
[0092] According to an embodiment of the present invention, the above weighted processing of the feature data to be processed by the triplet loss function by using the image resolution quality evaluation index as the weight of the triplet loss function to obtain the improved triplet loss function includes: calculating the Euclidean distance matrix between the feature data, and using the image resolution quality evaluation index to calculate the image resolution quality score ratio of the images in the multi-resolution image training dataset; preprocessing the image resolution quality score ratio, and obtaining the resolution weight by using the preprocessed image resolution quality score ratio; performing an operation on the resolution weight and the Euclidean distance matrix to obtain the Euclidean weighted distance matrix, and using the weighted distance matrix to obtain the farthest positive sample and the nearest negative sample of the images in the multi-resolution image training dataset; using the farthest positive sample, the nearest negative sample, and the Euclidean weighted distance matrix to improve the triplet loss function, so as to obtain the improved triplet loss function.
[0093] The following further details the improvement process of the above triplet loss function through specific embodiments.
[0094] The ratio of the high-frequency energy to the total energy , that is, the image resolution quality score is the ratio of the high-frequency energy to the total energy, multiplied by 1000, as shown in Equation (5):
[0095] (5).
[0096] Taking as the image resolution quality score (i.e., the image resolution quality evaluation index, the same below), constructing a triplet loss, and when calculating the Euclidean distance between the features in the triplet loss, using the corresponding image resolution quality score between the features Weight it as a weight:
[0097] Calculate the features of the Euclidean distance matrix , as shown in formula (6):
[0098] (6),
[0099] where and are the -th and -th features of the -th and -th samples respectively, and
[0100] Calculate the image resolution quality scoring ratio , as shown in formula (7):
[0101] (7).
[0102] If , take the reciprocal, as shown in formula (8):
[0103] (8).
[0104] Calculate the resolution weight , as shown in formula (9):
[0105] (9).
[0106] Multiply the distance matrix by the resolution weight to obtain the weighted distance matrix , as shown in formula (10):
[0107] (10),
[0108] where represents element-wise multiplication.
[0109] For each sample , find the farthest positive sample and the nearest negative sample , as shown in formulas (11) and (12):
[0110] (11),
[0111] (12),
[0112] where is a sample set of the same kind as the sample and is a sample set of a different kind from the sample .
[0113] Calculate the triplet loss , as shown in formula (13):
[0114] (13),
[0115] wherein is the number of samples is the margin of the triplet loss
[0116] According to an embodiment of the present invention, the above-mentioned feature extraction of the high-frequency component energy and the low-frequency component energy by using the pedestrian matching model respectively includes: performing wavelet decomposition on the low-frequency component energy by using the low-frequency identity flow network of the pedestrian matching model to obtain low-frequency component energy sub-bands; extracting global structure information related to identity from the low-frequency component energy sub-bands by using the low-frequency identity flow network to obtain low-frequency features; performing wavelet decomposition on the high-frequency component energy by using the high-frequency detail flow network of the pedestrian matching model to obtain high-frequency component energy sub-bands; extracting resolution-sensitive local detail features from the high-frequency component energy sub-bands by using the high-frequency detail flow network to obtain high-frequency features, thereby completing the high-frequency feature extraction and the low-frequency feature extraction based on wavelet-spatial domain two-stream feature decoupling
[0117] According to an embodiment of the present invention, the above-mentioned dynamic fusion of the extracted high-frequency features and low-frequency features to obtain two-stream fusion features includes: dynamically generating the coefficients of the resolution discriminator in the pedestrian matching model based on the image resolution quality evaluation index; dynamically fusing the high-frequency features and the low-frequency features by using the coefficients of the resolution discriminator based on the learnable gating mechanism to obtain two-stream fusion features
[0118] The following further details the process of obtaining the high-frequency and low-frequency features and the two-stream fusion features of the images in the multi-resolution image training dataset by the above-mentioned wavelet-spatial domain two-stream feature decoupling method through specific embodiments
[0119] During the training process of the pedestrian matching model, the wavelet-spatial domain two-stream feature decoupling method is adopted: using the low-frequency identity flow network (LL Stream) to process the low-frequency component energy to obtain low-frequency features; using the high-frequency detail flow network to extract the high-frequency component energy (LH / HL / HH sub-bands) to obtain high-frequency features; dynamically fusing the two-stream features through the learnable gating mechanism
[0120] Low-frequency Identity Stream Network (LL Stream): Perform wavelet decomposition on the input image, extract the low-frequency components (LL sub-band), and extract the global structure information related to identity through the low-frequency identity stream network to obtain low-frequency features 。
[0121] High-frequency Detail Stream Network (HH Stream): Perform wavelet decomposition on the input image, extract the high-frequency components (LH / HL / HH sub-bands), and extract the resolution-sensitive local detail features through the high-frequency detail stream network to obtain high-frequency features 。
[0122] Feature Fusion Module: Dynamically fuse the two-stream features through a learnable gating mechanism to obtain the final features ,as shown in Equation (14):
[0123] (14),
[0124] where is dynamically generated by the resolution classifier and adjusted according to the resolution quality score of the image of to achieve adaptive feature fusion for images with different resolutions.
[0125] Figure 3 is a flowchart of the cross-resolution pedestrian matching method according to an embodiment of the present invention.
[0126] As Figure 3 shown, the above cross-resolution pedestrian matching method includes operations S310 to S350.
[0127] In operation S310, use the trained pedestrian matching model to calculate the high and low frequency component energies of the query image set, and perform high and low frequency feature extraction based on wavelet-spatial domain two-stream feature decoupling on the high and low frequency component energies of the query image set to obtain the high and low frequency features of the query image set, where the trained pedestrian matching model is trained based on the training method of the above pedestrian matching model.
[0128] In operation S320, use the trained pedestrian matching model to dynamically fuse the high and low frequency features of the query image set based on a learnable gating mechanism to obtain the two-stream fusion features of the query image set.
[0129] In operation S330, use the trained pedestrian matching model to calculate the high and low frequency component energies of the reference image set, and perform high and low frequency feature extraction based on wavelet-spatial domain two-stream feature decoupling on the high and low frequency component energies of the reference image set to obtain the high and low frequency features of the reference image set, where the reference image set has a different resolution from the query image set.
[0130] In operation S340, the high-frequency and low-frequency features of the reference image set are dynamically fused based on a learnable gating mechanism using the trained pedestrian matching model to obtain the two-stream fusion features of the reference image set.
[0131] In operation S350, the trained pedestrian matching model is used to calculate the similarity between the two-stream fusion features of the query image set and the two-stream fusion features of the reference image set, and the pedestrians in the query image set and the reference image set are matched based on the similarity to obtain the pedestrian matching result.
[0132] The process of the cross-resolution pedestrian matching method will be further described in detail through specific embodiments below.
[0133] The query set (query image set) and the gallery set (reference image set) are input into the established recognition model to obtain the feature vectors of the cross-resolution images, and end-to-end real-time matching is performed to obtain the pedestrian matching result.
[0134] Perform wavelet decomposition on the input image, extract the low-frequency component (LL sub-band), and extract the global structure information related to the identity through the low-frequency identity flow network to obtain the low-frequency features , perform wavelet decomposition on the input image, extract the high-frequency components (LH / HL / HH sub-bands), and extract the resolution-sensitive local detail features through the high-frequency detail flow network to obtain the high-frequency features , dynamically fuse the two-stream features through the gating mechanism to obtain the final features , the features in the query set and the features in the gallery set are subjected to end-to-end real-time matching. By calculating the similarity between the feature vectors, it can be determined whether the pedestrians in the query set and the pedestrians in the gallery set are the same person.
[0135] Compared with the prior art, a cross-resolution pedestrian re-identification method based on wavelet transform improved triplet loss proposed by the present invention obtains an image resolution quality score through wavelet transform of the picture (image resolution quality evaluation index, the same below), and when calculating the Euclidean distance in the triplet loss function, the corresponding image resolution quality scores between the features are used as weights for weighted processing. Compared with the traditional triplet loss, the method of wavelet transform improved triplet loss takes into account the influence of resolution differences on feature representation, and through the image resolution quality score Reduce such influences and perform more stably when processing samples with different resolutions. At the same time, through the wavelet-spatial domain two-stream feature decoupling architecture, the identity and resolution features are explicitly separated to avoid the contamination of identity discrimination by resolution noise. While ensuring the inference speed, the inference accuracy and robustness are improved.
[0136] To verify the effectiveness of the present invention, in the campus learning situation monitoring system of a well-known university in Hefei, on the premise of fully informing the school, teachers and relevant students and obtaining the authorization of the school, teachers and relevant students, 1103 campus cameras were monitored. One frame of video image was collected and saved every 1 second, and real-time person re-identification was carried out. At present, the cross-resolution person re-identification technology with improved triplet loss proposed by the present invention has been deployed in the server and has been running smoothly for about one and a half years, and satisfactory results have been obtained.
[0137] Figure 4 It is a structural diagram of a cross-resolution person matching system according to an embodiment of the present invention.
[0138] As Figure 4 shown, the above-mentioned cross-resolution person matching system 400 includes a query set high-low frequency feature acquisition module 410, a query set two-stream fusion feature acquisition module 420, a reference set high-low frequency feature acquisition module 430, a reference set two-stream fusion feature acquisition module 440, and a similarity calculation and person matching module 450.
[0139] The query set high-low frequency feature acquisition module 410 is used to calculate the high-low frequency component energies of the query image set by using the trained person matching model, and perform high-low frequency feature extraction based on wavelet-spatial domain two-stream feature decoupling on the high-low frequency component energies of the query image set to obtain the high-low frequency features of the query image set, wherein the trained person matching model is trained based on the training method of the above-mentioned person matching model. In one embodiment, the query set high-low frequency feature acquisition module 410 can be used to perform the operation S310 described above, which will not be elaborated here.
[0140] The query set two-stream fusion feature acquisition module 420 is used to dynamically fuse the high-low frequency features of the query image set by using the trained person matching model based on a learnable gating mechanism to obtain the two-stream fusion features of the query image set. In one embodiment, the query set two-stream fusion feature acquisition module 420 can be used to perform the operation S320 described above, which will not be elaborated here.
[0141] The reference set high-low frequency feature acquisition module 430 is configured to calculate the high-low frequency component energies of the reference image set by using the trained pedestrian matching model, and perform high-low frequency feature extraction based on wavelet-spatial domain two-stream feature decoupling on the high-low frequency component energies of the reference image set to obtain the high-low frequency features of the reference image set, wherein the reference image set and the query image set have different resolutions. In one embodiment, the reference set high-low frequency feature acquisition module 430 may be configured to perform the operation S330 described above, which will not be elaborated herein.
[0142] The reference set two-stream fusion feature acquisition module 440 is configured to dynamically fuse the high-low frequency features of the reference image set based on a learnable gating mechanism by using the trained pedestrian matching model to obtain the two-stream fusion features of the reference image set. In one embodiment, the reference set two-stream fusion feature acquisition module 440 may be configured to perform the operation S340 described above, which will not be elaborated herein.
[0143] The similarity calculation and pedestrian matching module 450 is configured to calculate the similarity between the two-stream fusion features of the query image set and the two-stream fusion features of the reference image set by using the trained pedestrian matching model, and perform pedestrian matching on the query image set and the pedestrians in the reference image set based on the similarity to obtain the pedestrian matching result. In one embodiment, the similarity calculation and pedestrian matching module 450 may be configured to perform the operation S350 described above, which will not be elaborated herein.
[0144] According to an embodiment of the present invention, any combination of the query set high-low frequency feature acquisition module 410, the query set dual-stream fusion feature acquisition module 420, the reference set high-low frequency feature acquisition module 430, the reference set dual-stream fusion feature acquisition module 440, and the similarity calculation and pedestrian matching module 450 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the query set high-low frequency feature acquisition module 410, the query set dual-stream fusion feature acquisition module 420, the reference set high-low frequency feature acquisition module 430, the reference set dual-stream fusion feature acquisition module 440, and the similarity calculation and pedestrian matching module 450 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the query set high-low frequency feature acquisition module 410, the query set dual-stream fusion feature acquisition module 420, the reference set high-low frequency feature acquisition module 430, the reference set dual-stream fusion feature acquisition module 440, and the similarity calculation and pedestrian matching module 450 can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0145] Figure 5 It is a block diagram of an electronic device suitable for implementing the training method of the pedestrian matching model and the cross-resolution pedestrian matching method according to an embodiment of the present invention.
[0146] As Figure 5 shown, the electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage part 508 into the random access memory (RAM) 503. The processor 501 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 501 can also include on-board memory for caching purposes. The processor 501 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0147] In the RAM 503, various programs and data required for the operation of the electronic device 500 are stored. The processor 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. The processor 501 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 502 and / or the RAM 503. It should be noted that the programs may also be stored in one or more memories other than the ROM 502 and the RAM 503. The processor 501 may also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.
[0148] According to an embodiment of the present invention, the electronic device 500 may further include an input / output (I / O) interface 505, and the input / output (I / O) interface 505 is also connected to the bus 504. The electronic device 500 may further include one or more of the following components connected to the input / output (I / O) interface 505: an input portion 506 including a keyboard, a mouse, etc.; an output portion 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 508 including a hard disk, etc.; and a communication portion 509 including a network interface card such as a LAN card, a modem, etc. The communication portion 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read from it can be installed into the storage portion 508 as needed.
[0149] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0150] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 502 and / or RAM 503 and / or one or more memories other than ROM 502 and RAM 503.
[0151] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0152] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0153] The above describes the embodiments of the present invention. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A method for training a pedestrian matching model, characterized in that: The method comprises: The high-frequency component energy and the low-frequency component energy of the image in the multi-resolution image training data set are calculated using the pedestrian matching model, the total energy between the high-frequency component energy and the low-frequency component energy is calculated, and the high-frequency component energy is calculated with the total energy to obtain the image resolution quality evaluation index; Calculating the Euclidean distance matrix between the feature data, and using the image resolution quality evaluation index to calculate the image resolution quality score ratio of the images in the multi-resolution image training data set; Preprocessing the image resolution quality score ratio, and obtaining a resolution weight using the preprocessed image resolution quality score ratio; The resolution weight and the Euclidean distance matrix are operated to obtain a Euclidean weighted distance matrix, and the weighted distance matrix is used to obtain the farthest positive sample and the closest negative sample of the image in the multi-resolution image training data set; Improving the triplet loss function by using the farthest positive sample, the nearest negative sample and the Euclidean weighted distance matrix to obtain an improved triplet loss function; The pedestrian matching model is used to extract features of the high-frequency component energy and the low-frequency component energy respectively, and the extracted high-frequency features and low-frequency features are dynamically fused to obtain dual-stream fusion features; Using the improved triplet loss function to process the dual-stream fusion features and the true value labels corresponding to the dual-stream fusion features to obtain a triplet loss value, and using the triplet loss value to update the parameters of the pedestrian matching model; Iterate the feature extraction and dynamic fusion operations, triplet loss value calculation operations, and model parameter update operations until the preset training conditions are met to obtain a trained pedestrian matching model.
2. The method according to claim 1, characterized in that The high-frequency component energy and low-frequency component energy of the image in the multi-resolution image training dataset are calculated using the pedestrian matching model, including: Based on the authorization of the experimental user, the image of the experimental user is collected by setting the downsampling magnification of the image acquisition device to obtain the multi-resolution image training data set; Performing wavelet transform on the images in the multi-resolution image training data set by using the pedestrian matching model to obtain wavelet transform coefficients; An energy function is defined based on the wavelet transform coefficients, and the energy function is used to respectively calculate the high-frequency component energy of the high-resolution image and the low-frequency component energy of the low-resolution image in the multi-resolution image training data set.
3. The method according to claim 1, characterized in that Using the pedestrian matching model to extract features of the high-frequency component energy and the low-frequency component energy respectively includes: Using the low-frequency identity flow network of the pedestrian matching model, the low-frequency component energy is subjected to wavelet decomposition to obtain low-frequency component energy subbands; The low-frequency identity flow network is used to extract global structural information related to identity from the low-frequency component energy subband to obtain low-frequency features.
4. The method according to claim 3, characterized in that: Also includes: Using the high-frequency detail flow network of the pedestrian matching model, the high-frequency component energy is subjected to wavelet decomposition to obtain high-frequency component energy subbands; The high-frequency detail flow network is used to extract resolution-sensitive local detail features of the high-frequency component energy subband to obtain high-frequency features and then complete high-frequency feature extraction and low-frequency feature extraction based on wavelet-spatial domain dual-stream feature decoupling.
5. The method according to claim 1, characterized in that The extracted high-frequency features and low-frequency features are dynamically fused to obtain dual-stream fusion features including: Dynamically generate coefficients of a resolution classifier in the pedestrian matching model based on the image resolution quality evaluation index; Based on a learnable gating mechanism, the high-frequency features and the low-frequency features are dynamically fused using the coefficients of the resolution resolver to obtain the dual-stream fusion features.
6. A cross-resolution pedestrian matching method, characterized in that: The method comprises: The trained pedestrian matching model is used to calculate the high- and low-frequency component energies of the query image set, and the high- and low-frequency feature extraction based on wavelet-spatial dual-stream feature decoupling is performed on the high- and low-frequency component energies of the query image set to obtain the high- and low-frequency features of the query image set, wherein the trained pedestrian matching model is trained based on the method described in any one of claims 1 to 5; Using the trained pedestrian matching model, the high-frequency and low-frequency features of the query image set are dynamically fused based on a learnable gating mechanism to obtain dual-stream fusion features of the query image set; The trained pedestrian matching model is used to calculate the high- and low-frequency component energies of the reference image set, and high- and low-frequency feature extraction based on wavelet-spatial domain dual-stream feature decoupling is performed on the high- and low-frequency component energies of the reference image set to obtain high- and low-frequency features of the reference image set, wherein the reference image set and the query image set have different resolutions; Using the trained pedestrian matching model, the high-frequency and low-frequency features of the reference image set are dynamically fused based on a learnable gating mechanism to obtain dual-stream fusion features of the reference image set; The trained pedestrian matching model is used to calculate the similarity between the dual-stream fusion features of the query image set and the dual-stream fusion features of the reference image set, and the pedestrians in the query image set and the reference image set are matched based on the similarity to obtain a pedestrian matching result.
7. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Cross-resolution pedestrian re-identification method based on wavelet transform
CN114898410A