Pedestrian re-identification method, system and device for privacy protection

By combining high-frequency and low-frequency information in the federated learning framework, using DCT conversion and attention mechanisms, the balance between privacy protection and identification accuracy in the prior art is solved, and an efficient pedestrian re-identification method is achieved.

CN120071383AActive Publication Date: 2025-05-30BEIJING JIAOTONG UNIV

Patent Information

Application Number
CN202411891973.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-30
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

The existing anonymous pedestrian re-identification method based on frequency domain processing has led to a decrease in model identification accuracy while protecting privacy, especially due to the low utilization rate of low-frequency information, which leads to the lack of full utilization of all available information to optimize the accuracy of identification.

Method used

In the federated learning framework, combining high-frequency information and low-frequency information to reduce noise interference, extract and enhance pedestrian characteristic information through DCT conversion and attention mechanism, and segment data processing between the server and the client to ensure privacy protection and identification accuracy.

Benefits of technology

It realizes that while protecting pedestrian data privacy, maintaining a high recognition level of the model, reducing the risk of leakage after key cracking, and making full use of iterative supplements of low-frequency information, enriching the characteristics of high-frequency information in the identification task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071383A_ABST
    Figure CN120071383A_ABST
Patent Text Reader

Abstract

The invention relates to a pedestrian re-identification method for privacy protection. The pedestrian re-identification method comprises the following steps: S1, data preparation; S2, a preprocessing stage; s3, a frequency domain conversion stage: performing frequency domain conversion on the feature set obtained in the S2, representing a DCT coefficient in a frequency domain after frequency domain conversion by Xh and w, and segmenting a converted frequency domain picture into a key channel and a non-key channel; s4, a feature enhancement stage: S5, calculating the total loss of the cross entropy in the model pedestrian re-identification task and the total loss of the generator; s6, anonymized low-frequency information sent by the client is supplemented to the model, after the information is supplemented, the model carries out a pedestrian re-identification task again, and the cross entropy loss and the total loss of the generator are calculated; and S7, repeating the steps S4-S6 until the generated model test result accords with the expectation or reaches the number of training times.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of pedestrian recognition, and particularly to a pedestrian re-identification method, system and device with privacy protection. Background Art

[0002] As the pedestrian re-identification task is applied to more and more real-world scenarios (such as airports, passenger stations, etc.), a large amount of pedestrian data is input into the training of the recognition model. Pedestrian data provides rich feature information for training the model, which helps to improve the accuracy and robustness of the recognition algorithm. Moreover, by analyzing pedestrian data, the model can identify abnormal behavior patterns, prevent criminal activities, and enhance public safety. However, sensitive data containing pedestrian privacy is at risk of being leaked, which may lead to the abuse or public disclosure of personal sensitive information without authorization. To solve this problem, anonymous pedestrian re-identification, as a new solution, has received extensive attention from all walks of life in recent years. It encrypts pedestrian data to remove or blur personally identifiable information, such as facial features, to reduce the risk of privacy leakage, ensure security during storage and transmission, and achieve preliminary protection of real data. However, the blurring operation causes the pedestrian dataset to lack too much feature information, which may lead to the irreusability of the dataset and reduce the model's recognition ability. Anonymous pedestrian re-identification, privacy protection is not the end, and the ultimate goal is to maintain the original level of the model's recognition ability after privacy protection. Frequency domain processing, as one of the current advanced anonymization technologies, shows great potential in data privacy protection. This technology converts the original pedestrian image into a frequency domain image, extracts features that are not easily used to identify personal identity, such as gait and body shape contours, for re-identification tasks, and combines privacy protection algorithms such as differential privacy to add noise to the frequency domain image, reducing the risk of privacy leakage while maintaining the usability of the data.

[0003] Although the field of anonymous pedestrian re-identification has undergone years of research and development, current research work based on frequency domain processing mainly focuses on the extraction of high-frequency information. The model only extracts personally identifiable features that are not easily recognizable by the human eye, that is, high-frequency information. Although it reduces the risk of privacy leakage of pedestrian data, after discarding the low-frequency information, the model's attention cannot be concentrated on the effective visual features at key positions, which easily leads to a decrease in the model's recognition accuracy. For example, information such as faces, clothing, and body postures together constitute complete pedestrian data. However, after the model obtains the pedestrian contour, it cannot be directly applied to the recognition task because it cannot fully capture the specific features at these key positions, and the high-frequency noise will interfere with the model, which limits the model's ability to re-identify pedestrians.

[0004] In existing research methods for anonymous pedestrian re-identification based on frequency domain processing, the utilization rate of low-frequency information is relatively low. After many low-frequency information containing effective visual features are removed, they are not reasonably reused. These methods sacrifice the performance of the model to a certain extent, causing the model not to fully utilize all available information to optimize the recognition accuracy. Often, due to excessive focus on protecting privacy data, the impact on the accuracy of the recognition task is ignored. Existing methods mainly focus on the privacy protection role of high-frequency information but do not optimize in the possible high-frequency noise, which will interfere with the recognition ability of the model. And because high-frequency information lacks more pedestrian features, the re-identification ability using only high-frequency information for the recognition task is much lower than that of the original image. Existing methods do not apply strict privacy protection techniques to the pedestrian re-identification task. Common anonymous pedestrian re-identification techniques include key strategies. Due to their reversibility, once the key rule is cracked, there is a high risk of privacy leakage in the face of malicious attacks such as generation attacks. Summary of the Invention

[0005] The anonymous pedestrian re-identification method based on frequency domain processing of the present invention not only needs to protect the privacy data of pedestrians but also maintain a relatively high recognition level of the model. It not only adopts the frequency domain processing method but also applies a privacy protection algorithm. Excessive privacy protection usually causes a certain performance decline. Therefore, how to balance data privacy and model utility becomes an issue that cannot be ignored. Different from existing methods, we introduce a new method. By combining high-frequency information and low-frequency information in the federated learning framework, noise interference is reduced, and while improving the recognition accuracy of the model, more strict data privacy protection is provided for the pedestrian re-identification model. The specific technical solutions are as follows:

[0006] A privacy protection pedestrian re-identification method includes the following steps:

[0007] S1: Data preparation: Sample T frames of images from the video sequence of pedestrians as the video segment input to the model

[0008] S2: Preprocessing stage: Using the discreteness of DCT, use the backbone network of the convolutional neural network to extract the spatial features of each frame of image obtained in S1 to obtain the feature map, and the obtained feature set is represented as where C represents the number of channels, and H and W represent the height and width of the feature map respectively;

[0009] S3: Frequency domain conversion stage: X h,wX represents the DCT coefficients in the transformed frequency domain, and m(i, j) represents the coefficients of the feature map in the spatial domain, representing the rows and columns in the frequency domain respectively; that is, the value of the feature map m at the position (i, j) corresponds to the frequency response value (h, w) of the DCT spectrum X; and the transformed frequency domain image is segmented into key channels and non-key channels

[0010] S4: Feature enhancement stage: By calculating the L 2 norm in the channel dimension to construct the attention score, and then generating the attention map through normalization to highlight the most important regions or features in each frame of the video;

[0011] S5: Calculate the overall cross-entropy loss in the person re-identification task and the overall loss of the generator;

[0012] S6: Generator stage: The model will supplement the anonymized low-frequency information sent by the client. After supplementing the information, the model performs the person re-identification task again, and calculates its cross-entropy loss and the overall loss of the generator

[0013] S7: Repeat S4 - S6 until the model test results meet the expectations or reach the number of training times.

[0014] Optionally, the feature set in step S2 is calculated using the following formula

[0015]

[0016] where X h,w represents the DCT coefficients in the transformed frequency domain, and m(i, j) represents the coefficients of the feature map in the spatial domain, representing the rows and columns in the frequency domain respectively. That is, the value of the feature map m at the position (i, j) corresponds to the frequency response value (h, w) of the DCT spectrum X.

[0017] Optionally, step S3 specifically includes the following steps:

[0018] S31: Calculate the channel energy at the maximum response value X for each frequency band:

[0019]

[0020] where represents the c-th channel ability of the K-th frequency band, represents the DCT coefficients of the c-th channel

[0021] S32: Determine the threshold of threshold, according to C key ={c|E c >threshold}, the transformed frequency domain image is segmented into key channels and non-key channels, where the key channels contain low-frequency information and the non-key channels contain high-frequency information, Ckey Represents the set of channels selected as key channels.

[0022] S33: The high-frequency information is transmitted to the server side, and the low-frequency information is transmitted to the client side. The segmented image representation for model training and inference: Server side: Client side:

[0023] Optionally, the method for determining the threshold is as follows: First, according to the energy calculated for each frequency band, select the 5 channels with the highest energy from each of the Y, Cb, and Cr components as key channels, for a total of 15 channels. The remaining channels are regarded as non-key channels. Then, by calculating and analyzing the precision-recall curve of the re-identification task performed after the information of this group of classifications enters the server side. Calculate the precision and recall under different threshold groupings in this way, and finally select a threshold that maximizes the precision without reducing the recall.

[0024] Optionally, the step S4 specifically includes the following steps:

[0025] S41: Calculate the attention score of the high-frequency information frequency band entering the server side

[0026] S42: Aggregate the features of different frames and different frequency bands to extract the discriminative information of each frame:

[0027] Optionally, the step S4 further includes:

[0028] In the training stage, use the lowest frequency component in the frequency domain to enhance the representation of the shared features in the video sequence, focusing on extracting the features common to the entire video sequence to enhance the model's ability to identify pedestrian identities. By introducing frame-level attention mapping, the model can learn the importance of each frame in the video and use it to enhance the representation of the common features within the sequence:

[0029] A f = σ(Wight · m f + b)

[0030] where A f represents the attention mapping of the f-th frame, σ represents the sigmoid activation function, Wight and b represent the learned weights and biases respectively, and m f represents the feature map of the f-th frame.

[0031] Feature map of the entire sequence:

[0032]

[0033] Among them, m ′ represents the feature map of the entire video sequence, T represents the number of frames in the video sequence, and ⊙ represents element-by-element multiplication. During the training phase, the difference in attention maps between frames is used as a regularization term to bring the common information of frames in the video closer. The inter-frame regularization term is:

[0034]

[0035] Among them, R(m ′ ) represents the regularization term, which is used to reduce the difference in attention mapping between adjacent frames, ‖·‖ F represents the Frobenius norm.

[0036]

[0037] Among them, m shared represents the shared features of the entire video sequence, Indicates the lowest frequency component of the f-th frame.

[0038] Optionally, step S6 specifically includes the following steps:

[0039] S61: Extract feature maps of low-frequency information entering the client, average the channels, and generate feature masks: M mask =σ(W mask ·F client +b mask )

[0040] S62: Transmit the feature mask as supplementary information to the server to achieve feature transfer: F s ′ erver =F server +M norm ⊙F server

[0041] S63: The video image after supplementing the information is anonymized and updated based on GAN, and the generated dataset continues to be used for training and testing of the pedestrian re-identification task.

[0042] Optionally, the method for generating the anonymized updated image is:

[0043] Anonymized updated images are generated by learning a mapping function G:X→Y, where X is the updated pedestrian image set and Y is the anonymized image set. This mapping is learned through adversarial training, so that the generated anonymized images can further protect the privacy of pedestrians while retaining enough information for pedestrian re-identification (Re-ID) tasks;

[0044] The anonymization process involves a generator G X and a discriminator D Y . The generator G X is responsible for converting the original image into an anonymized image, and the discriminator D Y tries to distinguish between real images and the generated anonymized images.

[0045] A privacy-preserving person re-identification system, comprising:

[0046] A data preparation module for sampling T frames of images from a video sequence of a pedestrian as a video segment for model input

[0047] A preprocessing module for using the discreteness of DCT and performing spatial feature extraction on each frame of image obtained from S1 using a convolutional neural network backbone network to obtain a feature map, and the obtained feature set is represented as where C represents the number of channels, and H and W respectively represent the height and width of the feature map;

[0048] A frequency domain conversion module; for using X h,w to represent the DCT coefficients in the frequency domain after frequency domain conversion, and splitting the converted frequency domain image into key channels and non-key channels;

[0049] A feature enhancement module: for constructing an attention score by calculating the L 2 norm in the channel dimension, then generating an attention map through normalization to highlight the most important regions or features in each frame of the video; then calculating the overall cross-entropy loss in the person re-identification task and the overall loss of the generator;

[0050] A generator module: for supplementing the anonymized low-frequency information sent by the client, and after supplementing the information, the model performs the person re-identification task again, calculating its cross-entropy loss and the overall loss of the generator; continuously repeating the training until the model test result meets the expectation or reaches the number of training times.

[0051] An electronic device, characterized in that it includes a memory and a processor, and the memory is connected to the processor;

[0052] The memory stores computer instructions, and the processor executes the computer instructions to execute the above privacy-preserving person re-identification method.

[0053] A computer-readable storage medium, characterized in that the computer-readable storage medium stores computer instructions for causing a computer to execute the above privacy-preserving person re-identification method.

[0054] Compared with the prior art, the present invention has the following advantages:

[0055] Compared with the existing methods, the proposed frequency-domain processing-based anonymous pedestrian re-identification method realizes the protection of pedestrian data privacy in the federated learning framework without using additional public data and key policies, reduces the risk of leakage after key cracking, and makes full use of the iterative supplement of low-frequency information to enrich the features of high-frequency information in the recognition task. The anonymized dataset can still be used for the pedestrian re-identification task.

[0056] After the original image is transformed into a frequency-domain image, the high-frequency information channel and the low-frequency information channel are segmented. The high-frequency information enters the server side for the recognition task, and the low-frequency information enters the client side. The features that are easily recognizable by the human eye in the client are not directly transmitted to the server, realizing the privacy protection of specific features.

[0057] In order to balance data privacy protection and task applicability, the data that enters the server side and performs the recognition task focuses on reducing the features that are easily recognizable by the human eye. The pedestrian re-identification model learned by the attacker is difficult to recover and reconstruct because the data in the training process itself lacks many key features, avoiding the impact of generative attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the specific embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0059] Figure 1 It is a flowchart of the embodiment;

[0060] Figure 2 It is a schematic diagram of the model of the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail in conjunction with the drawings. The examples of the embodiments are shown in the drawings. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0062] Anonymous pedestrian re-identification based on frequency domain processing is a challenging privacy data protection task for pedestrian re-identification, aiming to maintain a high recognition level of the model while protecting the privacy data of pedestrians. The pedestrian re-identification (Re-ID) task analyzes pedestrian image or video data to identify and track the identity of specific individuals under different sensors (such as cameras), which plays an important role in public safety and monitoring efficiency, but also raises concerns about privacy protection. Current pedestrian re-identification models may involve personal privacy when training the dataset, and may lack transparency during the data collection process, where pedestrian data is captured and analyzed without the knowledge of the individuals. Since the model aims to complete the tasks of identifying and tracking individuals, access to the privacy data such as unauthorized pedestrian faces and postures hidden in the training and test datasets may disclose personal sensitive information. When facing some specific attacks, such as generative attacks, attackers may steal sensitive training data from the published trained re-identification model, resulting in serious privacy leakage problems. Moreover, current frequency domain processing research is mostly applied in face recognition tasks for operations such as feature extraction, image enhancement, and noise reduction. Therefore, the present invention proposes an effective method, namely an anonymous pedestrian re-identification method based on frequency domain processing, enabling the model to maintain a high recognition accuracy while protecting privacy. Specifically, we apply the federated learning framework. Through the discrete cosine transform network (DCT), the original pedestrian image is decomposed into a high-frequency channel image and a low-frequency channel image. The server-side processes the high-frequency channel representing abstract features, and the client-side processes the low-frequency channel representing specific features. For privacy protection, after removing the specific features, the pedestrian data on the server-side no longer has the specific features in human visual perception, effectively resisting malicious behaviors such as generative attacks. For the pedestrian re-identification task, we supplement some low-frequency information from the client-side in the server-side to improve the recognition accuracy.

[0063] The present invention proposes a new anonymous pedestrian re-identification model - a pedestrian re-identification framework with privacy preservation (Pedestrian Re-identification Framework with Privacy Preserving). Compared with existing technical solutions, it mainly has three innovations: The first is to introduce an attention-enhanced frequency selection module, which uses channel segmentation and attention mechanism in the server for processing high-frequency information to select representative frequencies of the pedestrian contour; the second is to add a content supplementation module based on feature mask generation. The high-frequency information is combined with partially anonymized low-frequency information for feature enhancement again to improve the client's judgment ability for pedestrian matching in the pedestrian re-identification task; the third is to combine a generative adversarial network in the client for processing low-frequency information to anonymize the low-frequency information. The anonymized data is used to supplement the server, providing privacy protection for the specific features of the original pedestrians and being able to better balance the accuracy and privacy of the model in the re-identification task. The following specifically explains the methods and modules designed and used in the present invention. The specific steps of this solution are as follows, and the flowchart is as Figure 1 shown, including:

[0064] S1: Data preparation: Given a video sequence of a pedestrian, assume that T frames of images are sampled as the video segment input to the model T represents the number of sampled frames

[0065] S2: In the preprocessing stage, using the discreteness of DCT, a convolutional neural network backbone network is used to extract spatial features from each frame of the image to obtain the feature map, and the obtained feature set can be expressed as where C represents the number of channels, and H and W represent the height and width of the feature map respectively

[0066] S3: In the frequency domain conversion stage, X h,w represents the DCT coefficient in the converted frequency domain, m(i,j) represents the coefficient of the feature map in the spatial domain, and represents the row and column in the frequency domain respectively. That is, the value of the feature map m at the position (i, j) corresponds to the frequency response value (h, w) of the DCT spectrum X. And the converted frequency domain image is segmented into key channels and non-key channels

[0067] S4: In the feature enhancement stage, calculate the L 2 norm in the channel dimension to construct the attention score, and then generate an attention map through normalization to highlight the most important regions or features in each frame of the video

[0068] S5: Calculate the overall cross-entropy loss in the pedestrian re-identification task and the overall loss of the generator

[0069] S6: In the generator stage, the model will supplement the anonymized low-frequency information sent by the client. After supplementing the information, the model performs the person re-identification task again, calculates its cross-entropy loss and the overall generator loss.

[0070] S7: Repeat S4 - S6 until the model test results meet the expectations or the training times are reached.

[0071] Traditional spatial domain methods may cause the destruction of spatial relationships due to spatial alignment problems between frames, which will affect the accuracy of pedestrian recognition. Compared with the method based on spatial segmentation, FSM uses DCT-based frequency information. In the frequency domain, information exists in the form of frequency components, and these components do not require precise alignment like pixels in the spatial domain. Therefore, even if there are minor spatial offsets or alignment problems between frames in the original video, their representations in the frequency domain can still maintain consistency, allowing the model to focus on those frequency components that are most useful for pedestrian recognition, rather than relying on spatial features that may vary due to alignment problems, effectively avoiding the problems of spatial alignment and destruction of spatial relationships, and improving the model's ability to understand video content. In the step S2, the DCT transformation includes:

[0072] The frequency selection module adopts the restricted random sampling method. Given a video sequence, assume that T frames of images are sampled as the video segment input to the model. T represents the number of sampled frames. Utilizing the discreteness of DCT, use the backbone network of convolutional neural network (CNN) to perform spatial features on each frame of image, extract the obtained feature maps, and then transform the feature maps in the spatial domain to the frequency domain. The feature set obtained through the CNN network can be expressed as where C represents the number of channels, and H and W represent the height and width of the feature map respectively. X is transformed to the frequency domain using DCT, and the corresponding DCT spectrum can be expressed as a specific formula:

[0073]

[0074] where, X h,w represents the DCT coefficient in the transformed frequency domain, m(i, j) represents the coefficient of the feature map in the spatial domain, and respectively represent the row and column in the frequency domain. That is, the value of the feature map m at the position (i, j) corresponds to the frequency response value (h, w) of the DCT spectrum X.

[0075] The step S3 specifically includes the following steps:

[0076] Frequency channel segmentation

[0077] According to the importance of frequency channels, they are divided into critical channels and non-critical channels. Critical channels contain most of the visual information, while non-critical channels are used for model training and inference on the server side to protect privacy.

[0078] To extract discriminative features of video frames in the frequency domain, most methods use global average pooling or horizontal spatial segmentation as frame-level representative features, which may have problems such as model information loss or inconsistent spatial alignment. This module retains effective information by using fine-grained frequency features instead of spatial features, increasing the effective information embedded in different frequency components, and can effectively avoid problems such as spatial alignment issues or spatial relationship disruption. Based on the assumption that the global average pooling result is one of the special frequency components in the DCT, the related formula is:

[0079]

[0080] where, and f(i, j) represent the DCT output and input respectively, H is the input dimension, and W is the applied frequency.

[0081] Formula (2) selects the frequency equal to 0, and the cosine term is equal to 1. The result shows that the average pooling result is equivalent to the lowest frequency component in the DCT (the component with frequency 0):

[0082]

[0083] X 00 represents the lowest frequency component in the DCT spectrum, that is, the global average pooling result.

[0084] According to the above proof, different DCT components can be applied to explore the most discriminative information in each frame, and the lowest DCT component is used to represent the shared information of the entire sequence. Specifically, the transformed frequency spectrum is divided into several frequency bands, and each frequency band is represented by its maximum response value, which can express effective information more compactly and help identify and extract the most discriminative features in each frame.

[0085] The entire frequency spectrum X is divided into K frequency bands, and each frequency band is represented by the maximum response value as follows:

[0086]

[0087] where, B k represents the range of the k-th frequency band.

[0088] Calculate the channel energy of each frequency band at the maximum response value X:

[0089]

[0090] Among them, it represents the c-th channel ability of the K-th frequency band. It represents the DCT coefficient of the c-th channel.

[0091] By selecting key channels and setting thresholds, it is used to determine which channels contain important visual information.

[0092] C key ={c|E c >threshold}

[0093] Among them, C key represents the set of channels selected as key channels, and threshold is a threshold used to determine which channels contain important visual information.

[0094] The method for determining the threshold of threshold is as follows: First, according to the energy calculated for each frequency band, select the 5 channels with the highest energy from each component of Y, Cb, and Cr as key channels, a total of 15 channels are selected, and the remaining channels are regarded as non-key channels. Then, by calculating and analyzing the precision-recall curve of the re-identification task executed after the information of this group of classifications enters the server side. Calculate the precision and recall under different threshold groupings in this way, and finally select a threshold that maximizes the precision without reducing the recall.

[0095] The high-frequency information transmitted to the server side and the low-frequency information transmitted to the client side are used for the segmentation image representation of model training and inference:

[0096]

[0097] Among them, I server and I client respectively represent the image representations used for model training and inference on the server side and the client side, C non-key represents the set of non-key channels, and C key represents the set of key channels.

[0098] 3) In step S4 described above, attention enhancement uses an attention map to learn the most discriminative information part in each frame. By calculating the L 2 norm in the channel dimension to construct an attention score, and then generating an attention map through normalization processing. This attention map highlights the most important regions or features in each frame of the video, thereby guiding the model to focus on the visual cues that are most helpful for the pedestrian re-identification task. The related formula is as follows:

[0099] a. Calculate the attention scores for each frequency band:

[0100]

[0101] where S c,x,y represents the attention score at channel c and position (x, y).

[0102] b. Generate a normalized attention map:

[0103]

[0104] where S c ′ ,x,y represents the normalized attention score.

[0105] c. Construct a 2D attention map:

[0106] M c,x,y = S c ′ ,x,y ·X x,y

[0107] where M c,x,y represents the weighted frequency response at channel c and position (x, y).

[0108] By aggregating features from different frames and different frequency bands, discriminative information for each frame is extracted while suppressing unimportant features. The extracted discriminative information can be expressed as:

[0109]

[0110] where D represents the aggregated discriminative feature.

[0111] In the training stage, this module uses the lowest frequency components in the frequency domain to enhance the representation of shared features in the video sequence, focusing on extracting features common to the entire video sequence to enhance the model's ability to identify pedestrian identities. By introducing frame-level attention mapping, the model can learn the importance of each frame in the video and use it to enhance the representation of shared features within the sequence:

[0112] A f = σ(Wight·m f + b)

[0113] where A f represents the attention mapping for the f-th frame, σ represents the sigmoid activation function, Wight and b represent the learned weights and biases respectively, and m f represents the feature map for the f-th frame.

[0114] Feature map of the entire sequence:

[0115]

[0116] Among them, m ′ represents the feature map of the entire video sequence, T represents the number of frames in the video sequence, and ⊙ represents element-by-element multiplication. During the training phase, the difference in attention maps between frames is used as a regularization term to bring the common information of frames in the video closer. The inter-frame regularization term is:

[0117]

[0118] Among them, R(m ′ ) represents the regularization term, which is used to reduce the difference in attention mapping between adjacent frames, ‖·‖ F represents the Frobenius norm.

[0119]

[0120] Among them, m shared represents the shared features of the entire video sequence, Indicates the lowest frequency component of the f-th frame.

[0121] The server-side model obtains the pedestrian's profile information through the frequency selection module, but because this information is visually unimportant high-frequency information, the server-side model accesses data, which may cause its attention to human features to be inaccurate, reducing its performance in pedestrian feature recognition. Therefore, an information interaction module (Interactive block) is introduced to allow attention to be transferred from the client to the server, and to compensate for the lack of visual information by passing information to the client model, thereby helping the server-side model to more accurately identify pedestrian features. The step S6 includes the following steps:

[0122] First, the client calculates the average value of each channel from the feature map and generates a single-channel feature mask. This mask reflects the importance of different regions in the feature map. Generation of feature mask:

[0123] M mask =σ(W mask ·F client +b mask )

[0124] Among them, M mask represents the feature mask, F client represents the feature map extracted by the client, W mask and b mask They represent the learnable weights and biases respectively, and σ represents the sigmoid activation function.

[0125] The client then normalizes the feature mask to the [0,1] interval:

[0126]

[0127] Among them, M norm represents the normalized feature mask, which is used to ensure that the value of the feature mask is within a reasonable range.

[0128] Transmit the normalized feature mask to the server side to achieve attention transfer:

[0129] F s ′ erver = F server + M norm ⊙ F server

[0130] Among them, F s ′ erver represents the updated server-side feature map, F server represents the original server-side feature map, and ⊙ represents element-wise multiplication, which is used to transfer the client's attention to the server side.

[0131] Finally, after the server side receives the feature mask, it is adjusted to be the same size as the feature map of the server-side model, then activated by a sigmoid function, and element-wise multiplied with the server-side feature map to update the server-side feature representation.

[0132] The video image after supplementary information is subjected to GAN-based anonymization generation, and the generated dataset continues to be used for training and testing of the person re-identification task. The updated representation image on the server side contains a part of the high-frequency information and low-frequency information of the original pedestrian data, and there may still be pedestrian sensitive information points. To avoid the interception of this part of information by attackers resulting in privacy leakage, it is necessary to anonymize the updated feature representation. The present invention uses conditional generative adversarial networks (cGANs) to implement the anonymization process of the updated image.

[0133] Anonymization of the updated image is generated by learning a mapping function G∶X→Y, where X is the set of updated pedestrian images and Y is the set of anonymized images. This mapping is learned through adversarial training, so that the generated anonymized images can further protect the privacy of pedestrians while retaining sufficient information for the person re-identification (Re-ID) task. The anonymization process involves a generator G X and a discriminator D Y . The generator G XResponsible for converting the original image into an anonymized image, the discriminator D Y Attempts to distinguish between real images and the generated anonymized images.

[0134] A computer device provided by an alternative embodiment of the present invention executes a privacy-preserving person re-identification method. The computer device includes: one or more processors, a memory, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface).

[0135] In some alternative embodiments, if needed, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system).

[0136] The processor can be a central processing unit, a network processor, or a combination thereof. Among them, the processor can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.

[0137] Among them, the memory stores instructions executable by at least one processor, so that the at least one processor executes the privacy-preserving person re-identification method shown in the above embodiments.

[0138] The memory can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required by at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory can include a high-speed random access memory and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory can include a memory remotely set relative to the processor, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0139] The memory may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid state drive; the memory may further include a combination of the above types of memory.

[0140] The computer device further includes an input device and an output device. The processor, the memory, the input device and the output device may be connected via a bus or other means.

[0141] The input device can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device, such as touch screen, keypad, mouse, trackpad, touchpad, pointing stick, one or more mouse buttons, trackball, joystick, etc. The output device may include a display device, an auxiliary lighting device (such as an LED), and a tactile feedback device (such as a vibration motor), etc. The above display device includes but is not limited to liquid crystal display, light emitting diode, display and plasma display. In some alternative embodiments, the display device may be a touch screen.

[0142] The computer device further includes a communication interface for the computer device to communicate with other devices or communication networks.

[0143] The embodiment of the present invention further provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading through a network the original computer code stored in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium, so that the method described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware.

[0144] Wherein, the storage medium may be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk or a solid state drive, etc.; further, the storage medium may further include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor or the hardware, the method shown in the above embodiments is implemented.

[0145] Embodiments of the present application may also provide a computer program product, including computer program instructions, which cause a processor to execute the steps in the above method when the computer program instructions are run by the processor. Among them, the computer program product can be written in any combination of one or more programming languages for executing the program code of the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language 10 or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0146] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.

Claims

1. A privacy-preserving person re-identification method, characterized in that: The steps include: S1: Data preparation: Sample T frames of images from the video sequence of pedestrians as video clips for model input S2: Preprocessing stage: Using the discreteness of DCT, the convolutional neural network backbone network is used to extract the spatial features of each frame image obtained in S1 to obtain the feature map. The obtained feature set is expressed as Where C represents the number of channels, H and W represent the height and width of the feature map respectively; S3: Frequency domain conversion stage: The feature set obtained in S2 is converted into frequency domain to X h,w Represent the DCT coefficients in the frequency domain after frequency domain conversion, and divide the converted frequency domain image into key channels and non-key channels; S4: Feature enhancement stage: The attention score is constructed by calculating the L2 norm of the channel dimension, and then the attention map is generated through normalization to highlight the most important areas or features in each frame of the video; Then train the generative model; S5: Calculate the overall cross entropy loss and the overall loss of the generator in the model pedestrian re-identification task; S6: Supplement the model with the anonymized low-frequency information sent by the client. After the supplemented information, the model once again performs the task of pedestrian re-identification and calculates its cross entropy loss and the overall loss of the generator; S7: Repeat S4-S6 until the generated model test results meet expectations or the number of training times is reached.

2. The privacy-preserving person re-identification method according to claim 1, characterized in that: The feature set in step S3 is calculated using the following formula: h∈{0,1,…,H-1},w∈{0,1,…,W-1} Among them, X h,w represents the DCT coefficient in the frequency domain after conversion, m(i, j) represents the coefficient of the feature map in the spatial domain, and represents the row and column in the frequency domain respectively; that is, the value of the feature map m at position (i, j) corresponds to the frequency response value (h, w) of the DCT spectrum X.

3. The privacy-preserving person re-identification method according to claim 1, characterized in that: The step S3 specifically comprises the following steps: S31: Calculate the channel energy of each frequency band at the maximum response value X: Where, represents the c-th channel capability of the K-th frequency band, Represents the DCT coefficient of the cth channel; S32: Determine the threshold value of threshold, and then key ={c|E c >threshold} rule, the converted frequency domain image is divided into key channels and non-key channels, where the key channels contain low-frequency information and the non-key channels contain high-frequency information. key represents the set of channels selected as key channels; S33: High-frequency information is transmitted to the server, and low-frequency information is transmitted to the client. Segmented image representation for model training and reasoning: Server: Client:

4. The privacy-preserving person re-identification method according to claim 1, characterized in that: The step S4 specifically comprises the following steps: S41: Calculate the attention score of the high-frequency information frequency band entering the server S42: Aggregate features of different frames and frequency bands to extract discriminative information for each frame:

5. The privacy-preserving person re-identification method according to claim 1, characterized in that: The step S4 further comprises: During the training phase, the lowest frequency component in the frequency domain is used to enhance the representation of shared features in the video sequence, focusing on extracting common features in the entire video sequence to enhance the model's ability to identify pedestrians. By introducing frame-level attention mapping, the model is able to learn the importance of each frame in the video and use it to enhance the representation of common features within the sequence: A f =σ(Wight·m f +b) Among them, A f represents the attention map of the f-th frame, σ represents the sigmoid activation function, Wight and b represent the learned weights and biases respectively, and m f represents the feature map of the f-th frame; Feature graph of the entire sequence: Among them, m ′ represents the feature map of the entire video sequence, T represents the number of frames in the video sequence, and ⊙ represents element-by-element multiplication. During the training phase, the difference in attention maps between frames is used as a regularization term to bring the common information of frames in the video closer. The inter-frame regularization term is: Among them, R(m ′ ) represents the regularization term, which is used to reduce the difference in attention mapping between adjacent frames, ‖·‖ F represents the Frobenius norm; Among them, m shared represents the shared features of the entire video sequence, Indicates the lowest frequency component of the f-th frame.

6. The privacy-preserving person re-identification method according to claim 1, characterized in that: The step S6 specifically comprises the following steps: S61: Extract feature maps of low-frequency information entering the client, average the channels, and generate feature masks: M mask =σ(W mask ·F client +b mask ); S62: Transmit the feature mask as supplementary information to the server to achieve feature transfer: F s ′ erver =F server +M norm ⊙F server ; S63: The video image after supplementing the information is anonymized and updated based on GAN, and the generated dataset continues to be used for training and testing of the pedestrian re-identification task.

7. The privacy-preserving person re-identification method according to claim 1, characterized in that: The method for generating the anonymized updated image is: Anonymized updated images are generated by learning a mapping function G:X→Y, where X is the updated pedestrian image set and Y is the anonymized image set. This mapping is learned through adversarial training, so that the generated anonymized images can further protect the privacy of pedestrians while retaining enough information for pedestrian re-identification (Re-ID) tasks; The anonymization process involves a generator G X and a discriminator D Y . Generator G X Responsible for converting the original image into an anonymized image, the discriminator D Y Try to distinguish between real images and generated anonymized images.

8. A privacy-preserving person re-identification system, characterized in that: include: The data preparation module is used to sample T frames of images from the video sequence of pedestrians as video clips for model input The preprocessing module is used to utilize the discreteness of DCT and use the convolutional neural network backbone network to extract the spatial features of each frame image obtained by S1 to obtain the feature map. The obtained feature set is expressed as Where C represents the number of channels, H and W represent the height and width of the feature map respectively; Frequency domain conversion module; used to h,w Represent the DCT coefficients in the frequency domain after frequency domain conversion, and divide the converted frequency domain image into key channels and non-key channels; Feature enhancement module: constructs an attention score by calculating the L2 norm of the channel dimension, and then generates an attention map through normalization to highlight the most important areas or features in each frame of the video; Then train the generative model; then calculate the overall cross entropy loss and the overall loss of the generator in the pedestrian re-identification task; Supplementary information module: used to supplement the anonymized low-frequency information sent by the client. After the supplementary information, the model once again performs the task of pedestrian re-identification, calculates its cross entropy loss and the overall loss of the generator; and uses the feature enhancement module and the supplementary information module to repeat the training until the model test results meet expectations or the training times are reached.

9. An electronic device, characterized in that: comprising a memory and a processor, wherein the memory and the processor are connected; The memory stores computer instructions, and the processor executes the privacy-preserving pedestrian re-identification method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the privacy-preserving pedestrian re-identification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for estimating timing error under OFDM system

    CN104092640A

  • Pedestrian re-identification method and device, equipment and medium

    CN111881757A

  • Network training method and device, pedestrian re-identification method and device, electronic equipment and storage medium

    CN112001321A

Cited By

  • Personnel re-identification method based on space-frequency feature selection and background consistency constraint

    CN122200750A