A Pedestrian Tracking Method Based on Federated Learning and Edge Computing

Through the method of combining federated learning and edge computing, cross-modal processing and optimization training technology are adopted to solve the data security and identification accuracy problems of pedestrian tracking technology in different scenarios, and realize cross-modal pedestrian tracking and privacy protection.

CN114582011BActive Publication Date: 2025-07-18GUANGXI PUBLIC INFORMATION IND CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111612745.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-07-18
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

The existing pedestrian tracking technology has data security and privacy issues, and cannot adapt to different scenarios, especially at night scenarios, and lacks cross-modal recognition capabilities.

Method used

Using a method of combining federated learning and edge computing, cross-modal processing technology and knowledge distillation and weight allocation optimization model training is used to achieve cross-modal pedestrian tracking, protect data privacy and improve model adaptability.

Benefits of technology

Cross-modal pedestrian tracking is realized to protect data security, solve the problems of degradation of identification accuracy and model convergence in multiple scenarios, and ensure user privacy and data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114582011B_ABST
    Figure CN114582011B_ABST
Patent Text Reader

Abstract

The present invention discloses a pedestrian tracking method based on federated learning and edge computing, comprising the following steps: Step 1, detecting and acquiring pedestrian images; Step 2, performing cross-modal processing on the pedestrian images; Step 3, extracting pedestrian features; Step 4, matching features to determine the target pedestrian. The model training based on federated learning in this method can effectively ensure data privacy and solve the data silo problem; the method of knowledge distillation and weight allocation is used to solve the problems of accuracy decline and model convergence caused by data heterogeneity in multiple scenarios; cross-modal processing technology is used to process images of different modalities to meet the requirements of cross-modal pedestrian tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision and federated learning, and particularly relates to a pedestrian tracking method based on federated learning and edge computing. Background Art

[0002] Current pedestrian tracking methods need to use a large amount of pedestrian data to train models. However, a single dataset often only contains data of a specific scenario. When the model trained based on such a dataset is applied to the actual scenario, the recognition accuracy often drops significantly. If data collected by video acquisition devices in multiple different scenarios is used for model training, it will lead to data leakage, and user privacy and data security cannot be guaranteed. In addition, existing pedestrian tracking technologies can only be applied to RGB video images taken during the day. However, pedestrian tracking not only needs to meet the use during the day but also needs to be applicable to night scenes. To sum up, the models obtained by training with a single labeled dataset in existing pedestrian tracking technologies have the disadvantages of poor generalization ability, inability to adapt to different usage scenarios, and inability to process cross-modal data. Therefore, a pedestrian tracking method that can ensure data security, adapt to different application scenarios, and achieve cross-modal recognition is the key to solving this problem.

[0003] As a distributed machine learning technology, federated learning enables multiple participating parties to perform machine learning while protecting data privacy and meeting the requirements of legality and compliance, thus solving the data silo problem. Edge computing technology is mainly applied in the field of video processing and has the advantages of low transmission latency, large bandwidth, and high computing performance. Applying federated learning and edge computing to the field of pedestrian tracking and using edge servers and cloud servers jointly for model training based on federated learning can effectively ensure data privacy and solve the data silo problem.

[0004] Chinese Patent CN201910119437.X discloses a pedestrian tracking method, device, and equipment. The method includes: obtaining a video frame to be detected; detecting candidate pedestrians in the video frame to be detected; extracting candidate pedestrian features of the candidate pedestrians; determining the difference between the candidate pedestrian features and the features already saved in the feature queue, and when the difference meets a preset condition, determining the candidate pedestrian as the target pedestrian; where the features already saved in the feature queue are features matching the target pedestrian. It can improve the efficiency of the pedestrian tracking process. However, this method can only meet the use during the day and cannot be applied to night scenes.

[0005] Chinese Invention Patent CN201510548633.0 discloses a video pedestrian detection and tracking method based on motion information and trajectory association. Pedestrian detection: Use the frame difference method to detect motion, and combine it with the morphological method in digital image processing. First, detect the motion area in the video, then extract features by adopting a sliding window search method in the motion area, and use a pre-trained pedestrian detection classifier to finally obtain the classification result. Tracking method: Use the pedestrian detection result obtained in the previous step as the input of this step. At the beginning, initialize a tracker for each detected pedestrian. Each tracker contains the historical motion information and appearance information of the target. When processing the current frame, for each detection result input, extract the position information and appearance information, and based on this, establish an association matrix to associate the tracking targets in the previous frame, and finally obtain the tracking trajectory of the pedestrian. The present invention has good real-time performance and also has good robustness in relatively complex scenarios. However, this method has a data security problem. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention provides a pedestrian tracking method based on federated learning and edge computing. This method uses an edge server and a cloud server to jointly perform model training based on federated learning, which can effectively ensure data privacy and solve the data island problem; uses the method of knowledge distillation and weight allocation to solve the problems of accuracy decline and model convergence caused by data heterogeneity in multiple scenarios when applying federated learning to pedestrian tracking; uses cross-modal processing technology to process images of different modalities to meet the needs of cross-modal pedestrian tracking.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A pedestrian tracking method based on federated learning and edge computing, comprising the following steps:

[0009] Step 1, detect and obtain pedestrian images;

[0010] Step 2, perform cross-modal processing on the pedestrian images;

[0011] Step 3, extract pedestrian features;

[0012] Step 4, match features to determine the target pedestrian.

[0013] The further description of the present invention is that the specific steps for detecting and obtaining pedestrian images in Step 1 include:

[0014] 11) Process the video to be detected to obtain several video frame images;

[0015] 12) Input each video frame image of the video file into the pedestrian tracking network. Determine whether a pedestrian appears in the current frame image through the pedestrian tracking network. If a pedestrian is detected, obtain and output the pedestrian image;

[0016] The detection and acquisition of the pedestrian image are based on the detection framework of PP-YOLOv2.

[0017] Further description of the present invention, the specific process of the cross-modal processing in step two is as follows:

[0018] 21) Judge the image data type: If it is an RGB image, execute step 22); if it is an IR image, execute step 23);

[0019] 22) Convert the RGB image to a grayscale image. At this time, the image is single-channel. Create a second channel and fill the second channel with 0 values;

[0020] 23) Use the IR image itself as the second channel and fill the first channel with 0 values;

[0021] 24) The image data processed through step 22) or step 23) is uniformly in a two-channel structure and output to step three for pedestrian feature extraction.

[0022] During the process of using the training set for model training, different nodes will selectively deactivate different modalities in the two-channel structure, so as to achieve the effect of cross-modal re-identification.

[0023] Further description of the present invention, the specific process of extracting pedestrian features in step three is as follows:

[0024] 31) Input the pedestrian image to be queried and the pedestrian images detected from the video and processed through cross-modal processing into the trained feature extraction model based on federated learning and edge computing respectively;

[0025] 32) The feature extraction model extracts the feature vectors of the pedestrian image to be queried and each detected pedestrian image respectively.

[0026] The model training method of the federated learning and edge computing is as follows. First, the cloud server selects K out of all N edge servers to participate in the model training to obtain an initial global model, and the cloud server distributes the global model to each edge server. After receiving the global model, the edge server connects the global model with the local classifier obtained in the previous round of training to form a new local model, and uses the local dataset to perform local update training on the model. The edge server saves the local classifier and uploads the parameters of the backbone network to the cloud server. Finally, the cloud server aggregates the model parameters uploaded by all the edge servers it receives to obtain a new global model, and continues to distribute the global model, repeating the above steps until the model converges and meets the accuracy requirement.

[0027] For further description of the present invention, the specific process of matching features and determining the target pedestrian in step 4 is as follows:

[0028] 41) Compare the feature vectors of the image to be queried with the feature vectors of the pedestrian images detected in the video respectively, and calculate the feature similarity; the formula for calculating the feature similarity is:

[0029]

[0030]

[0031] where x is the feature of the pedestrian to be queried, y is the feature of the pedestrian image detected in the video image, n is the feature dimension, and T is the cosine value between the two features; x,y is the cosine value between the two features;

[0032] 42) Take the pedestrians in the detected images that meet the similarity threshold as the target pedestrians in the video.

[0033] For further description of the present invention, the cloud server, edge server and data acquisition device are required in the training process of the feature extraction model; when the edge server and the cloud server transfer information, the data is encrypted using the homomorphic encryption algorithm before transmission.

[0034] For further description of the present invention, the cloud server, edge server and data acquisition device are required in the training process of the feature extraction model; when the edge server and the cloud server transfer information, the data is encrypted using the homomorphic encryption algorithm before transmission.

[0035] For further description of the present invention, the training method of the feature extraction model is optimized through knowledge distillation and weight allocation.

[0036] ​For further description of the present invention, the optimization method of knowledge distillation is as follows: First, a dataset is used as a shared dataset, and the cloud server distributes the shared dataset together with the initialized model parameters to the edge servers. Each edge server uses the local model trained with local data to predict the shared dataset and obtains predicted labels. Then, the updated weights of the local model and the obtained predicted labels are uploaded. The cloud server performs weighted averaging on the received predicted labels. Finally, the cloud server uses the shared dataset and the predicted labels to train the federated model, reducing the instability of the model and enabling the model to converge better.

[0037] For further description of the present invention, the optimization method of weight allocation is as follows: Using the cosine distance as a metric, a larger weight is assigned to the local model with greater parameter changes in each training, so that more newly learned knowledge is reflected in the global model. The specific process is as follows:

[0038] 51) The edge server randomly selects a set of training data D from the local training data batch ;

[0039] 52) In the next round of training, the edge server receives the model parameters sent by the cloud server, connects the global model with the local classifier obtained in the previous round of training to form a new local model, and saves the local model parameters and D batch as a log

[0040] 53) The edge server continues to train to obtain a new local model and saves the new model parameters and D batch as a log

[0041] 54) The edge server calculates the weight by averaging the cosine distances of each data point in D batch . The calculation formula is as follows: The calculation formula is as follows:

[0042]

[0043] 55) The edge server uploads the weight to the cloud server, and the cloud server uses it as the weight value for weighted averaging of each model.

[0044] The present invention has the following beneficial effects:

[0045] 1. The present invention processes images of different modalities by using cross-modal processing technology to meet the requirements of cross-modal pedestrian tracking.

[0046] 2. By using the method of knowledge distillation and weight redistribution, the present invention solves the problems of accuracy degradation and model convergence caused by data heterogeneity in multiple scenarios when applying federated learning to pedestrian tracking.

[0047] 3. The present invention conducts distributed training of the pedestrian tracking model through federated learning and edge servers, completing the entire process without data leaving the local area, protecting the security of monitoring data and user privacy. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a flow chart of pedestrian tracking.

[0049] Figure 2 It is a working flow chart of the cross-modal processing module.

[0050] Figure 3 It is a schematic diagram of the training process of the pedestrian feature extraction module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The present invention will be further described below in conjunction with the accompanying drawings.

[0052] Embodiment 1:

[0053] A pedestrian tracking method based on federated learning and edge computing includes the following steps:

[0054] Step 1, detect and obtain a pedestrian image;

[0055] Step 2, perform cross-modal processing on the pedestrian image;

[0056] Step 3, extract pedestrian features;

[0057] Step 4, match the features to determine the target pedestrian.

[0058] Embodiment 2:

[0059] A pedestrian tracking method based on federated learning and edge computing includes the following steps:

[0060] Step 1, detect and obtain a pedestrian image; specifically:

[0061] 11) Process the video to be detected to obtain a number of video frame images;

[0062] 12) Input each video frame image of the video file into the pedestrian tracking network, and determine whether a pedestrian appears in the current frame image through the pedestrian tracking network. If a pedestrian is detected, obtain and output the pedestrian image;

[0063] Step 2, perform cross-modal processing on the pedestrian image;

[0064] The specific process of the cross-modal processing is as follows:

[0065] 21) Determine the image data type: If it is an RGB image, execute step 22); if it is an IR image, execute step 23).

[0066] 22) Convert the RGB image to a grayscale image. At this time, the image is single-channel. Create a second channel and fill the second channel with 0 values.

[0067] 23) Use the IR image itself as the second channel and fill the first channel with 0 values.

[0068] 24) The image data processed through step 22) or step 23) is uniformly in a two-channel structure and output to step three for pedestrian feature extraction.

[0069] Step three, extract pedestrian features; the specific process is as follows:

[0070] 31) Input the pedestrian image to be queried and the pedestrian images detected from the video and processed through cross-modal respectively into the trained feature extraction model based on federated learning and edge computing.

[0071] 32) The feature extraction model extracts the feature vectors of the pedestrian image to be queried and each detected pedestrian image respectively.

[0072] Step four, match features and determine the target pedestrian; the specific process is as follows:

[0073] 41) Compare the feature vector of the image to be queried with the feature vectors of the pedestrian images detected in the video, and calculate the feature similarity; the formula for calculating the feature similarity is:

[0074]

[0075]

[0076]

[0077] x,y where x is the feature of the pedestrian to be queried, y is the feature of the pedestrian image detected in the video image, n is the feature dimension, and T

[0078] 42) Regard the pedestrians in the detected images that meet the similarity threshold as the target pedestrians in the video.

[0078] For further description of the present invention, a cloud server, an edge server and a data acquisition device are required during the training process of the feature extraction model; when the edge server transmits information to the cloud server, the data is encrypted using a homomorphic encryption algorithm before transmission.

[0079] Example 3:

[0080] A pedestrian tracking method based on federated learning and edge computing, comprising the following steps:

[0081] Step 1, detecting and obtaining pedestrian images; specifically including:

[0082] 11) Processing the video to be detected to obtain a number of video frame images;

[0083] 12) Inputting each video frame image of the video file into the pedestrian tracking network, and determining whether a pedestrian appears in the current frame image through the pedestrian tracking network. If a pedestrian is detected, obtain and output the pedestrian image;

[0084] Step 2, performing cross-modal processing on the pedestrian images;

[0085] The specific process of the cross-modal processing is as follows:

[0086] 21) Judging the image data type: if it is an RGB image, execute step 22); if it is an IR image, execute step 23);

[0087] 22) Converting the RGB image to a grayscale image. At this time, the image is single-channel, create a second channel and fill the second channel with 0 values;

[0088] 23) Using the IR image itself as the second channel and filling the first channel with 0 values;

[0089] 24) The image data processed through step 22) or step 23) is uniformly in a two-channel structure and output to step 3 for pedestrian feature extraction.

[0090] Step 3, extracting pedestrian features; the specific process is as follows:

[0091] 31) Inputting the pedestrian image to be queried and the pedestrian images detected from the video and processed through cross-modal processing into the trained feature extraction model based on federated learning and edge computing respectively;

[0092] 32) The feature extraction model extracts the feature vectors of the pedestrian image to be queried and each detected pedestrian image respectively.

[0093] During the training process of the feature extraction model, a cloud server, an edge server and data acquisition devices are required; when the edge server transmits information to the cloud server, the data is encrypted using a homomorphic encryption algorithm before transmission.

[0094] The training method of the feature extraction model is optimized through knowledge distillation and weight allocation;

[0095] The optimization method of the knowledge distillation is as follows: First, a dataset is used as the shared dataset, and the cloud server distributes the shared dataset together with the initialized model parameters to the edge servers. Each edge server uses the local model trained with local data to predict the shared dataset and obtains the predicted labels. Then, the updated weights of the local model and the obtained predicted labels are uploaded. The cloud server performs weighted averaging on the received predicted labels. Finally, the cloud server uses the shared dataset and the predicted labels to train the federated model, reducing the instability of the model and enabling the model to converge better;

[0096] The optimization method of the weight allocation is as follows: Using the cosine distance as the metric standard, a larger weight is assigned to the local model with greater parameter changes in each training, so that more newly learned knowledge is reflected in the global model. The specific process is as follows:

[0097] 51) The edge server randomly selects a set of training data D from the local training data batch ;

[0098] 52) In the next round of training, the edge server receives the model parameters sent by the cloud server, connects the global model with the local classifier obtained in the previous round of training to form a new local model, and saves the local model parameters and D batch as a log

[0099] 53) The edge server continues to train to obtain a new local model Save the new model parameters and D batch as a log

[0100] 54) The edge server calculates the weight by averaging the cosine distances of each data point in D batch The calculation formula is as follows: The calculation formula is as follows:

[0101]

[0102] 55) The edge server uploads the weight to the cloud server, and the cloud server uses it as the weight value for weighted averaging of each model.

[0103] Step Four, match features and determine the target pedestrian. The specific process is as follows:

[0104] 41) Compare the feature vector of the image to be queried with the feature vectors of the pedestrian images detected in the video respectively, and calculate

[0105] the feature similarity. The calculation formula of the feature similarity is as follows:

[0106]

[0107] Among them, x is the feature of the pedestrian to be queried, y is the feature of the pedestrian image detected in the video image, n is the feature dimension, and T x,y is the cosine value between the two features;

[0108] 42) The pedestrians in the detected images that meet the similarity threshold are used as the target pedestrians in the video.

[0109] The above embodiments are only exemplary embodiments of the present invention and are not used to limit the present invention. The protection scope of the present invention is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present invention, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present invention.

Claims

1. A pedestrian tracking method based on federated learning and edge computing, characterized in that It includes the following steps: Step 1, detect and obtain a pedestrian image; Step 2, perform cross-modal processing on the pedestrian image; The specific process of the cross-modal processing is as follows: 21) Judge the image data type: if it is an RGB image, execute step 22); if it is an IR image, execute step 23); 22) Convert the RGB image into a grayscale image. At this time, the image is single-channel. Create a second channel and fill the second channel with 0 values; 23) Use the IR image itself as the second channel and fill the first channel with 0 values; 24) The image data after being processed by step 22) or step 23) is uniformly in a two-channel structure and output to step 3 for pedestrian feature extraction; Step 3, extract pedestrian features. The specific process is as follows: 31) Input the pedestrian image to be queried and the pedestrian images detected from the video and processed by cross-modal processing into the trained feature extraction model based on federated learning and edge computing respectively; 32) The feature extraction model extracts the feature vectors of the pedestrian image to be queried and each detected pedestrian image respectively; Step 4, match features to determine the target pedestrian. The specific process is as follows: 41) Compare the feature vectors of the image to be queried with the feature vectors of the pedestrian images detected in the video respectively, and calculate the feature similarity; the formula for calculating the feature similarity is: where x is the feature of the pedestrian to be queried, y is the feature of the pedestrian image detected in the video image, n is the feature dimension, and T x,y is the cosine value between the two features; 42) Take the pedestrians in the detected images that meet the similarity threshold as the target pedestrians in the video.

2. The pedestrian tracking method based on federated learning and edge computing according to claim 1, characterized in that: The detection and acquisition of the pedestrian image in step 1 specifically includes: 11) Process the video to be detected to obtain several video frame images; 12) Input each video frame image of the video file into the pedestrian tracking network. Determine whether there is a pedestrian in the current frame image through the pedestrian tracking network. If a pedestrian is detected, obtain and output the pedestrian image.

3. The pedestrian tracking method based on federated learning and edge computing according to claim 2, wherein: The cloud server, edge server and data acquisition device are required during the training process of the feature extraction model; When the edge server transmits information to the cloud server, the data is encrypted using the homomorphic encryption algorithm before transmission.

4. The pedestrian tracking method based on federated learning and edge computing according to claim 3, characterized in that: The training method of the feature extraction model is optimized through knowledge distillation and weight allocation.

5. The pedestrian tracking method based on federated learning and edge computing according to claim 4, characterized in that: The optimization method of the knowledge distillation is as follows: First, use a dataset as the shared dataset. The cloud server distributes the shared dataset and the initialized model parameters to the edge servers together. Each edge server uses the local model trained with local data to predict the shared dataset to obtain prediction labels; then upload the updated weights of the local model and the obtained prediction labels. The cloud server performs weighted averaging on the received prediction labels; finally, the cloud server uses the shared dataset and the prediction labels to train the federated model, reduce the instability of the model, and make the model converge better.

6. The pedestrian tracking method based on federated learning and edge computing according to claim 5, characterized in that: The optimization method of the weight allocation is as follows: Use the cosine distance as the metric standard. For the local model with greater parameter changes in each training, allocate a greater weight to reflect more newly learned knowledge in the global model; the specific process is as follows: 51) The edge server randomly selects a set of training data D from the local training data batch ; 52) In the next round of training, the edge server receives the model parameters sent by the cloud server, connects the global model with the local classifier obtained in the previous round of training to form a new local model, and saves the local model parameters batch associated with D 53) The edge server continues training to obtain a new local model The new model parameters and D batch are saved as a log 54) The edge server calculates the weight by averaging the cosine distances of each data point in D batch The calculation formula is as follows: The calculation formula is as follows: 55) The edge server uploads the weights to the cloud server, and the cloud server uses them as the weight values for weighted averaging of each model.

Citation Information

Patent Citations

  • Video Pedestrian Detection and Tracking Method Based on Motion Information and Trajectory Association

    CN105224912B

  • A pedestrian tracking method, device and equipment

    CN109919043B

  • Cross-platform multi-modal public opinion analysis method based on federal learning and edge calculation

    CN113642700A