A Privacy Protection Method for Multimodal Pedestrian Behavior Monitoring
Through a differential privacy protection method of spatiotemporal hybrid Mamba and adaptive feature density perception, combined with dual-flow dynamic Transformer and dynamic learning rate, the privacy protection problems of sparseness and multimodal pedestrian behavior data are solved, and the accuracy and effectiveness of monitoring are improved.
Patent Information
- Application Number
- CN202510481109.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The prior art has the risk of privacy leakage when processing sparse and multimodal pedestrian behavior data, and existing privacy protection methods will lead to reduced data availability and poor model training results.
Multimodal pedestrian data fusion with space-time hybrid Mamba is adopted, combined with differential privacy protection of adaptive feature density perception, and pre-training is used to use dual-flow dynamic Transformer to protect sparse gradient privacy for adaptive space-time correlation, and highly available multimodal pedestrian behavior monitoring is carried out through dynamic learning rates.
While protecting pedestrian privacy, it improves the accuracy and effectiveness of multimodal pedestrian behavior monitoring, ensuring data availability and model robustness.
Smart Images

Figure CN120012164B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of information security and multimodal pedestrian behavior monitoring, and particularly relates to a privacy protection method for multimodal pedestrian behavior monitoring. Background Art
[0002] With the acceleration of the urbanization process and the popularization of intelligent devices, pedestrian behavior monitoring, as an important part of intelligent pedestrians and urban management, has gradually received extensive attention. The pedestrian behavior monitoring technology based on the cloud-edge-terminal collaborative architecture can efficiently process and analyze massive trajectory data by using distributed computing resources, predict the movement behaviors of individuals or groups in the future time period, and provide strong support for applications such as urban pedestrian flow management, public security monitoring, and intelligent navigation. Especially in scenarios such as autonomous driving, intelligent transportation, and emergency response, efficient and accurate pedestrian behavior monitoring can significantly improve the response speed and decision-making accuracy of the intelligent pedestrian system.
[0003] Although the pedestrian behavior monitoring technology based on cloud-edge-terminal collaboration has played an important role in improving the efficiency and reliability of the intelligent pedestrian system, in reality, pedestrian behavior data often exhibits sparsity and multimodality. Especially in complex urban scenarios, pedestrian behavior data usually comes from different sensors and terminal devices, and has different acquisition frequencies, spatial distributions, and data qualities. These data often contain sensitive personal information, such as trajectory routes, staying positions, etc. If directly uploaded to untrusted edge nodes, attackers can use means such as link attacks, man-in-the-middle attacks, and side-channel attacks to obtain personal sensitive information in the user's trajectory, resulting in user privacy leakage.
[0004] In the prior art, to protect the privacy of pedestrians between terminal devices and edge nodes, researchers have proposed some privacy protection technologies, including encryption methods, anonymization techniques, differential privacy, etc. However, encryption methods usually require a large amount of computing resources. Especially when encrypting and decrypting large-scale data, it may lead to a decrease in the system processing speed and response time. A core problem of anonymization techniques is the risk of re-identification. Even if the data has been anonymized, attackers may still be able to recover the true identity of the anonymized data through various methods (such as association analysis, external data, etc.). The core method of differential privacy is to blur the query results by injecting noise to protect the privacy of individuals. However, existing differential privacy methods are mainly applied to dense data. Facing highly sparse pedestrian behavior data, they will add a large amount of noise, resulting in a significant reduction in data availability, thus affecting the training effect and prediction accuracy of the model. Especially in the pedestrian behavior monitoring task, the monitoring ability of the model usually depends on a large amount of historical pedestrian behavior data and multimodal information. When these data become sparse after being scrambled, the model may not be able to fully mine the regularities in the data, thereby reducing its robustness and accuracy in practical applications. Summary of the Invention
[0005] This invention is made in view of the above problems, and its purpose is to provide a privacy protection method for multi-modal pedestrian behavior monitoring. A multi-modal pedestrian data fusion module based on spatio-temporal hybrid Mamba performs spatio-temporal multi-modal fusion on local multi-modal pedestrian user data; a differential privacy protection module based on adaptive feature density perception dynamically allocates privacy budgets to protect the fused data of multi-modal pedestrian users; a multi-modal pedestrian behavior pre-training module based on dual-flow dynamic Transformer filters redundant and irrelevant noise features and realizes pre-training for multi-modal pedestrian behavior monitoring; for sparse gradient privacy protection facing adaptive spatio-temporal correlation, calculates spatio-temporal correlation to dynamically adjust the size of Top-K to achieve gradient sparsification and protection of gradient spatio-temporal correlation; a highly available multi-modal pedestrian behavior monitoring method based on dynamic learning rate adjusts weights with dynamic learning rate and distributes them to edge nodes to achieve highly accurate pedestrian behavior monitoring. Based on the nature of differential privacy, it satisfies differential privacy and improves the security and effectiveness of multi-modal pedestrian behavior monitoring.
[0006] Specifically, the first aspect of this invention provides a privacy protection method for multi-modal pedestrian behavior monitoring, including the following steps:
[0007] Step 1: A multi-modal pedestrian data fusion method based on spatio-temporal hybrid Mamba extracts the feature of each modal pedestrian data and performs efficient spatio-temporal feature fusion for multi-modal.
[0008] Mamba is based on the Selective State Space Model (SSM) and realizes input-related information selection by dynamically adjusting parameters. Compared with the quadratic complexity of Transformer, Mamba has a linear time complexity and is especially suitable for processing long sequence data (such as video frames, time series signals).
[0009] Multi-modal pedestrian data refers to a pedestrian data set containing various types of information, including video, image, and numerical modalities.
[0010] Step 2: A differential privacy protection method based on adaptive feature density perception perturbs the fused features with Laplace noise to protect the correlation privacy of multi-modal spatio-temporal features;
[0011] Step 3: A multi-modal pedestrian behavior pre-training method based on dual-flow dynamic Transformer balances sparse noise features and dense noise features to improve the pre-training accuracy of the model;
[0012] Step 4: A sparse gradient privacy protection method facing adaptive spatio-temporal correlation protects the spatio-temporal correlation privacy of the model gradient;
[0013] Step 5: Based on the high-availability multi-modal pedestrian behavior monitoring method with dynamic learning rate, dynamically update the global model parameters to achieve highly accurate pedestrian behavior monitoring.
[0014] Further, the first step includes the following steps:
[0015] Step 1.1: Given the multi-modal pedestrian data of the user, extract the spatio-temporal features of each modality through the video encoder and the numerical encoder respectively;
[0016] Both the video encoder and the numerical encoder are single-modal encoders, and the formulas are as follows:
[0017] Video encoder:
[0018] ;
[0019] ;
[0020] Where: is the video feature of the nth layer after time-domain dependence modeling;
[0021] is the video feature of the nth layer;
[0022] is the video feature of the (n - 1)th layer;
[0023] is the layer normalization function;
[0024] is the time-domain dependence modeling module;
[0025] is the feed-forward layer;
[0026] Numerical encoder:
[0027] ;
[0028] ;
[0029] ;
[0030] ;
[0031] Where: is the output of the update gate;
[0032] is the sigmoid activation function;
[0033] is the weight matrix of the update gate;
[0034] is the hidden state of the GRU network at the previous moment;
[0035] is the attention weight vector;
[0036] is the Hadamard product;
[0037] is the input of the GRU network;
[0038] is the output of the reset gate;
[0039] is the weight matrix of the reset gate;
[0040] is the weight matrix of the candidate hidden state;
[0041] is the candidate hidden state;
[0042] is the hyperbolic tangent function;
[0043] is the hidden state of the GRU network at the current moment;
[0044] Step 1.2: Through the visual encoder combined with the time-domain dependence modeling module, the numerical encoder introduces a temporal dynamic gating mechanism to extract the spatio-temporal correlation between modalities;
[0045] The formula of the time-domain dependence modeling module is as follows:
[0046] ;
[0047] ;
[0048] Among them: is the visual feature of the g-th group and the n-th layer;
[0049] is the activation function;
[0050] is the learnable time relationship parameter;
[0051] is the visual feature of the (n - 1)-th layer;
[0052] is the time-domain dependence modeling module;
[0053] is the linear transformation function;
[0054] is a splicing function;
[0055] is the visual feature of the nth layer in the first group;
[0056] The formula of the time-series dynamic gating mechanism is as follows:
[0057] ;
[0058] Where: is the attention weight vector;
[0059] is the activation function;
[0060] is the scoring weight matrix;
[0061] is the hidden state of the previous moment of the GRU network;
[0062] is the input of the GRU network;
[0063] is the dimension for the input value;
[0064] Step 1.3: Based on the extracted spatio-temporal features and spatio-temporal correlations, perform cross-modal feature fusion through the blending module and the Mamba module, generate hybrid features using vector-level multiplication and vector-level addition operations, and capture spatio-temporal correlations through the efficient spatial scanning 2D layer, and finally output the multi-modal fusion pedestrian features.
[0065] The formula of the blending module is as follows:
[0066] ;
[0067] ;
[0068] ;
[0069] ;
[0070] ;
[0071] ;
[0072] Where: is the visual feature embedding of the ith layer obtained after passing through the self-attention layer;
[0073] is the layer normalization function;
[0074] is the self-attention layer;
[0075] is the visual embedding of the (i - 1)-th layer;
[0076] is the numerical feature embedding of the i-th layer obtained after passing through the self-attention layer;
[0077] is the numerical embedding of the (i - 1)-th layer;
[0078] is the similarity with the visual representation of the previous layer;
[0079] is the L2 norm, defined as the square root of the sum of the squares of the elements of the vector;
[0080] is the similarity with the numerical representation of the previous layer;
[0081] is the visual feature embedding of the i-th layer;
[0082] is the feed-forward layer;
[0083] is the numerical feature embedding of the i-th layer;
[0084] When the similarity with the visual representation of the previous layer or the similarity with the numerical representation of the previous layer is greater than the set threshold, the feature embedding of the previous layer is reused.
[0085] The Mamba module uses vector-level multiplication and vector-level addition operations to generate hybrid features and captures spatio-temporal correlations through an efficient spatial scan 2D layer, and finally outputs multi-modal fusion pedestrian features. The formula is as follows:
[0086] ;
[0087] ;
[0088] ;
[0089] ;
[0090] ;
[0091] Where: is the hybrid feature;
[0092] is the depth convolution operation;
[0093] is a linear transformation layer;
[0094] is a visual feature embedding;
[0095] is an element-wise multiplication operation;
[0096] is an element-wise addition operation;
[0097] is a numerical feature embedding;
[0098] is a visual mixed feature embedding;
[0099] is a layer normalization function;
[0100] is an efficient spatial scan 2D layer;
[0101] is a numerical mixed feature embedding;
[0102] is a final mixed feature embedding;
[0103] is a multi-modal fusion pedestrian feature;
[0104] is a channel attention module;
[0105] Furthermore, step two includes the following steps:
[0106] Step 2.1: Calculate the grid trajectory density based on the multi-modal fusion pedestrian flow features. The formula is as follows:
[0107] ;
[0108] Where: is the trajectory density of grid i;
[0109] is the number of features of grid i;
[0110] is the total sum of the number of features in the entire grid area;
[0111] Dynamically adjust the grid resolution through the regional feature distribution. By iteratively dividing the high-density areas, ensure the refined representation of sparse and dense areas. The formula is as follows:
[0112] ;
[0113] Where: is the resolution vector of grid i;
[0114] is the area size of grid i;
[0115] is the trajectory density of grid i;
[0116] Step 2.2: Dynamically allocate the privacy budget for each iteration based on the grid resolution and the initial privacy budget, calculate the remaining budget, and combine the regional trajectory density weights to re-allocate the remaining privacy budget;
[0117] Dynamically allocate the privacy budget for each iteration based on the grid resolution and the initial privacy budget. The formula is as follows:
[0118] ;
[0119] ;
[0120] ;
[0121] where: is the maximum number of iterations;
[0122] is the adjustment factor;
[0123] is the division size of the sub-unit;
[0124] is the multi-modal fusion pedestrian feature;
[0125] is the modulo operation on the multi-modal fusion pedestrian feature;
[0126] is the trajectory density of grid i;
[0127] is the privacy budget for the c-th iteration;
[0128] is the number of iterations;
[0129] is the predefined privacy budget;
[0130] is the privacy budget for the i-th iteration;
[0131] is the resolution vector of grid i;
[0132] Calculate the remaining privacy budget, and combine the regional trajectory density weights to re-allocate the remaining privacy budget. The formula is as follows:
[0133] Calculate the remaining privacy budget:
[0134] ;
[0135] Calculate the regional trajectory density weight:
[0136] ;
[0137] Quadratically allocate the remaining privacy budget:
[0138] ;
[0139] Where: is the remaining privacy budget;
[0140] is the predefined privacy budget;
[0141] is the total number of grids;
[0142] is the privacy budget for the i-th iteration;
[0143] is the regional trajectory density weight of grid i;
[0144] is the resolution vector of grid i;
[0145] is the trajectory density of grid i;
[0146] is the privacy budget for quadratic allocation;
[0147] Step 2.3: Add Laplace noise to the multi-modal fusion pedestrian features of each grid cell to generate the perturbed features. By adjusting and balancing the noise intensity based on the grid resolution and the grid trajectory feature density, ensure data availability. The formula is as follows:
[0148] ;
[0149] Where: is the perturbed feature of grid i;
[0150] is the multi-modal fusion pedestrian feature of grid i;
[0151] is used to balance the grid-based noise added;
[0152] is the Laplace noise;
[0153] is the privacy budget for secondary allocation;
[0154] is for balancing the added region feature-based noise;
[0155] is the privacy budget for the i-th iteration;
[0156] Furthermore, step three includes the following steps:
[0157] Step 3.1: Design a sparse self-attention branch based on squared LeakyReLU. Filter high query-key matching score features through non-linear activation, filter out low-correlation trajectory features, avoid redundant calculations, and at the same time use the negative value retention property of LeakyReLU to alleviate the neuron death problem. The formula is as follows:
[0158] ;
[0159] Where: is the sparse self-attention branch based on squared LeakyReLU;
[0160] is the activation function;
[0161] is the query matrix;
[0162] is the transpose of the key matrix;
[0163] is the feature dimension;
[0164] Step 3.2: Construct a standard dense self-attention branch. Use Softmax to comprehensively capture potential key features in multi-modal pedestrians and prevent over-dilution of sparse region features. The formula is as follows:
[0165] ;
[0166] Where: is the standard dense self-attention branch;
[0167] is the activation function;
[0168] is the query matrix;
[0169] is the transpose of the key matrix;
[0170] is the feature dimension;
[0171] Step 3.3: Propose a dynamic weight fusion mechanism to adaptively adjust the dual-branch attention weights through learnable parameters, balance the noise suppression in dense regions and the feature enhancement in sparse regions. The formula is as follows:
[0172] ;
[0173] ;
[0174] Where: is the nth learnable parameter;
[0175] is the nth learnable adjustment parameter;
[0176] is the modulo operation on the resolution vector of grid i;
[0177] is the resolution vector of grid i;
[0178] is the total number of grids;
[0179] is the dual-branch attention weight;
[0180] is the sparse self-attention branch based on squared LeakyReLU;
[0181] is the standard dense self-attention branch;
[0182] is the 1st learnable parameter;
[0183] is the 2nd learnable parameter;
[0184] is the value matrix;
[0185] Step 3.4: Integrate the social Transformer decoder, encode the neighbor trajectory feature matrix into social interaction embeddings through linear transformation, and extract pedestrian social constraints (such as avoidance and parallel behavior) through the encoder-decoder attention layer to generate trajectory features that conform to social rationality. The formula is as follows:
[0186] ;
[0187] Where: is the input embedding of the decoder;
[0188] is the linear embedding;
[0189] For pedestrian social constraint feature embedding;
[0190] Is a learnable parameter matrix;
[0191] Step 3.5: Design a dual monitoring mechanism, combine a regression monitoring head and a classification monitoring head (outputting pedestrian behavior and corresponding probabilities), construct a joint loss function based on the nearest neighbor clustering center, and jointly optimize pedestrian behavior Error and modal cross-entropy loss, generate the pre-training model gradient, and the formula is as follows:
[0192] ;
[0193] Where: Is the pre-training model gradient;
[0194] Is the first parameter of the balanced loss function;
[0195] Is Loss function;
[0196] Is the true pedestrian behavior;
[0197] Is the predicted pedestrian behavior;
[0198] Is the second parameter of the balanced loss function;
[0199] Is the cross-entropy loss function;
[0200] Is the true pedestrian behavior Corresponding probability;
[0201] Is the predicted pedestrian behavior Corresponding probability;
[0202] Furthermore, the fourth step includes the following steps:
[0203] Step 4.1: Select a subset in the cloud server, calculate the original gradient calculated by each edge node using the sample s in the t-th round of training, and correct the original gradient through error compensation to obtain the corrected gradient. The formula is as follows:
[0204] ;
[0205] ;
[0206] ;
[0207] Wherein: is the original gradient calculated by edge node m using sample s in the t-th round of training;
[0208] is the derivative of the loss function;
[0209] are the model parameters of edge node m in the t-th round;
[0210] is the number of samples;
[0211] is the error compensation of edge node m in the t-th round of training;
[0212] is the average gradient of edge node m in the (t - 1)-th round of training;
[0213] is the original gradient of edge node m in the (t - 1)-th round of training;
[0214] is the corrected gradient;
[0215] is the adjustment parameter used to control the influence of error compensation on the gradient;
[0216] is the maximum value function;
[0217] is the L2 norm, defined as the square root of the sum of the squares of the elements of the vector;
[0218] is the clipping threshold;
[0219] Step 4.2: Calculate the average gradient of each edge node, and calculate the spatio-temporal correlation based on the average gradient of the previous round, and dynamically adjust the size of Top-K. The formula is as follows:
[0220] ;
[0221] ;
[0222] ;
[0223] Wherein: is the average gradient of edge node m in the t-th round;
[0224] is the mini-batch data sample is the number of samples of;
[0225] is a small batch of data samples;
[0226] is the corrected gradient;
[0227] is the similarity between the average gradient of the t-th round and the average gradient of the (t - 1)-th round;
[0228] is the average gradient of the edge node m in the (t - 1)-th round;
[0229] is the L2 norm, defined as the square root of the sum of the squares of the elements of the vector;
[0230] is the Top-K size;
[0231] is the sparsification operation;
[0232] is the modulo operation on the average gradient vector of the edge node m in the (t - 1)-th round;
[0233] Step 4.3: Sparsify the average gradient of each edge node through the Top-K operator to obtain the sparse gradient, and record the gradient difference as the new error compensation. The formula is as follows:
[0234] ;
[0235] ;
[0236] ;
[0237] Where: is the average gradient of the edge node m in the t-th round obtained after the similarity operation;
[0238] is the similarity between the average gradient of the t-th round and the average gradient of the (t - 1)-th round;
[0239] is the average gradient of the edge node m;
[0240] is the sparse gradient of the edge node m;
[0241] is the sparsification operation;
[0242] is the new error compensation of the edge node m;
[0243] is the gradient difference;
[0244] Step 4.4: For each component of the sparse gradient, the edge node samples two variables, and then quantizes the gradient by the quantization step size. The edge node sends the quantized gradient to the cloud server. The formula is as follows:
[0245] ;
[0246] ;
[0247] ;
[0248] where: is the quantization step size;
[0249] is the first sampled variable, ;
[0250] is the component of the sparse gradient;
[0251] is the standard deviation of the Gaussian noise to be simulated;
[0252] is the maximum function;
[0253] is a hyperparameter used to limit the lower bound of the gradient value;
[0254] is the th component of the sparse gradient of edge node m in the t-th round;
[0255] is the L2 norm, defined as the square root of the sum of the squares of the elements of the vector;
[0256] is the quantized gradient of edge node m;
[0257] is the second sampled variable, ;
[0258] is the quantization function used to map continuous gradient values to discrete values;
[0259] is the floor function;
[0260] is the parameter of the function. In the second formula, x is ;
[0261] Furthermore, the fifth step includes the following steps:
[0262] Step 5.1: After the cloud server receives the quantized gradients from all edge nodes, it samples the same variables using the shared random seed of edge node m, and then decodes the gradients. The formula is as follows:
[0263] ;
[0264] Where: is the decoded quantized gradient of edge node m;
[0265] is the quantized gradient of edge node m;
[0266] is the second sampled variable, ;
[0267] Step 5.2: The cloud server aggregates all the decoded gradients, adjusts the weights according to the contribution degree of each edge node, and updates the global model using the dynamic learning rate, and starts the next iteration. The formula is as follows:
[0268] ;
[0269] ;
[0270] ;
[0271] Where: is the global gradient;
[0272] is the number of edge nodes participating in the t-th round of training;
[0273] is the model parameter of edge node m in the t-th round;
[0274] is the decoded gradient;
[0275] is the moving average of the squared gradient at the t-th iteration;
[0276] is the exponential decay rate of the squared gradient at the t-th iteration;
[0277] is the moving average of the squared gradient at the (t - 1)-th iteration;
[0278] is the learning rate at the t-th iteration;
[0279] is a constant used to prevent division by zero error;
[0280] is the global model parameter for the (t + 1)-th round;
[0281] is the global model parameter for the t-th round;
[0282] Step 5.3: The cloud server distributes the global model parameters to each edge node to perform the multi-modal pedestrian behavior monitoring task. In the model prediction stage, the model outputs multiple pedestrian behaviors and uses the behavior diversity enhancement algorithm to optimize the selection strategy to cover the multi-modal pedestrian behavior patterns. The formula is as follows:
[0283] ;
[0284] ;
[0285] where: is the similarity between the pedestrian behavior numbered i and the pedestrian behavior numbered j;
[0286] is the time step;
[0287] is the total time step;
[0288] is the pedestrian behavior numbered i at time step ;
[0289] is the pedestrian behavior numbered j at time step ;
[0290] is the maximization operation;
[0291] is the total number of pedestrians;
[0292] is the probability function;
[0293] is the probability corresponding to the behavior of the pedestrian numbered i;
[0294] is the similarity between the pedestrian behavior numbered i and the pedestrian behavior numbered j, and the total number of pedestrians is K;
[0295] is the adjustment parameter used to adjust the diversity of pedestrian behaviors;
[0296] is the pedestrian behavior numbered i;
[0297] is the pedestrian behavior numbered j. Brief Description of the Drawings
[0298] In order to more clearly illustrate the technical solutions in the embodiments of the present drawings or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present drawings. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0299] Figure 1 is the flowchart of the steps of the present invention;
[0300] Figure 2 is a comparison chart of data availability between the embodiments of the present invention and traditional privacy protection methods for pedestrian behavior monitoring under the total privacy budget of different multi-modal fusion data;
[0301] Figure 3 is a comparison chart of data availability between the embodiments of the present invention and traditional privacy protection methods for pedestrian behavior monitoring under the privacy budget for protecting the gradients of different edge models;
[0302] Figure 4 is a comparison chart of the average distance error between the embodiments of the present invention and traditional privacy protection methods for pedestrian behavior monitoring under different training batches;
[0303] Figure 5 is a comparison chart of the final distance error between the embodiments of the present invention and traditional privacy protection methods for pedestrian behavior monitoring under different training batches.
[0304] The realization, functional features, and advantages of the objectives of the present drawings will be further described in conjunction with the embodiments with reference to the drawings. Detailed Embodiments
[0305] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following will describe and explain the present invention in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments provided by the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0306] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, the present invention can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in such a development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present invention, some design, manufacturing, or production changes based on the technical content disclosed in the present invention are only conventional technical means and should not be understood as insufficient disclosure of the content of the present invention.
[0307] If there is no special instruction, all embodiments and optional embodiments of the present invention can be combined with each other to form new technical solutions.
[0308] If there is no special instruction, all technical features and optional technical features of the present invention can be combined with each other to form new technical solutions.
[0309] If there is no special instruction, all steps of the present invention can be carried out in sequence or randomly, preferably in sequence. For example, the method includes steps (a) and (b), which means that the method can include steps (a) and (b) carried out in sequence, or can also include steps (b) and (a) carried out in sequence. For example, it is mentioned that the method may further include step (c), which means that step (c) can be added to the method in any order. For example, the method can include steps (a), (b), and (c), or can also include steps (a), (c), and (b), or can also include steps (c), (a), and (b), etc.
[0310] If there is no special instruction, the "including" and "comprising" mentioned in the present invention mean open-ended or can also be closed-ended. For example, the "including" and "comprising" can mean that other components not listed can also be included or comprised, or can also only include or comprise the listed components.
[0311] If there is no special instruction, in the present invention, the term "or" is inclusive. For example, the phrase "A or B" means "A, B, or both A and B". More specifically, any of the following conditions satisfies the condition "A or B": A is true (or exists) and B is false (or does not exist); A is false (or does not exist) while B is true (or exists); or both A and B are true (or exist).
[0312] To better understand the solutions of the embodiments of the present invention, some related terms and concepts that may be involved in the embodiments of the present invention will be introduced below.
[0313] (1) Artificial intelligence (AI), also known as intelligent machinery or machine intelligence, refers to machines created by humans that can exhibit intelligence. Generally, artificial intelligence refers to the technology that presents human intelligence through ordinary computer programs.
[0314] (2) Deep learning (DL), deep learning is a method of machine learning. Its concept stems from the research of artificial neural networks. A multi-layer perceptron with multiple hidden layers is a deep learning structure, so deep learning is often referred to as a deep neural network. Compared with general machine learning, deep learning can automatically extract features, that is, automatically combine simple features into more complex features and use these combinations for multi-layer weight learning to solve problems. The motivation for studying deep learning is to build a neural network that simulates the human brain for analysis and learning, which mimics the mechanism of the human brain to interpret data, such as images, sounds, and texts. Deep learning first emerged in image recognition, but in just a few years, deep learning has been extended to various fields of machine learning and has shown excellent performance. It has applications in various major fields such as image recognition, speech recognition, audio processing, natural language recognition, robot bioinformatics processing, search engines, human-computer gaming, online advertising targeted delivery, medical automatic diagnosis, and finance.
[0315] (3) A single-modal encoder is a neural network module that specifically processes a single data modality (such as text, image, audio). Its core goal is to extract features and perform representation learning on the input data without involving cross-modal interactions. For example, in a visual question answering (VQA) task, a single-modal encoder will separately use a language encoder (to process text) and a visual encoder (to process images), extract intra-modal features through the self-attention mechanism (Self-Attention) and a feed-forward network (FFN), and then fuse them through a cross-modal encoder.
[0316] (4) Mamba is a sequence modeling architecture based on the state space model (SSM). Its core innovations include: dynamically adjusting parameters (such as input-related matrices B, C, Δ), deciding which information to retain or ignore based on the current input, solving the static parameter limitation of traditional SSMs; leveraging the GPU memory hierarchy, reducing memory IO overhead through scanning rather than convolution calculations, achieving linear time complexity and a 5-fold increase in inference throughput; and not requiring the traditional attention mechanism, only combining SSM with multi-layer perceptron (MLP) blocks to form a more concise structure.
[0317] (5) The Transformer is a deep learning model architecture used for natural language processing (NLP) and other sequence-to-sequence tasks. It was first proposed by Vaswani et al. in 2017. The Transformer architecture introduced the self-attention mechanism, which is a key innovation and enables it to perform well in processing sequence data.
[0318] (6) Differential Privacy is a data protection technology that ensures that the modification of a single data record does not significantly affect the output result by introducing random noise into the algorithm, thereby protecting user privacy. It is commonly used in scenarios such as federated learning or training with sensitive data.
[0319] In this embodiment, as Figure 1 shown, a privacy protection method for multi-modal pedestrian behavior monitoring includes the following steps:
[0320] Step 1: Based on the spatio-temporal hybrid Mamba multi-modal pedestrian data fusion method, extract the features of pedestrian data in each modality and perform efficient fusion of multi-modal spatio-temporal features;
[0321] Mamba is based on the selective state space model (SSM) and realizes input-related information selection by dynamically adjusting parameters. Compared with the quadratic complexity of the Transformer, Mamba has a linear time complexity and is particularly suitable for processing long sequence data (such as video frames, time series signals).
[0322] Multi-modal pedestrian data refers to a pedestrian data set containing various types of information, including video, image, and numerical modalities.
[0323] Step 2: Based on the differential privacy protection method of adaptive feature density perception, perform Laplace noise perturbation on the fused features to protect the correlation privacy of multi-modal spatio-temporal features;
[0324] Step 3: Based on the multi-modal pedestrian behavior pre-training method of the dual-flow dynamic Transformer, balance sparse noise features and dense noise features to improve the pre-training accuracy of the model;
[0325] Step 4: A sparse gradient privacy protection method for adaptive spatio-temporal correlation to protect the spatio-temporal correlation privacy of the model gradient;
[0326] Step 5: A highly available multi-modal pedestrian behavior monitoring method based on dynamic learning rate to dynamically update the global model parameters and achieve highly accurate pedestrian behavior monitoring.
[0327] Furthermore, Step 1 includes the following steps:
[0328] Step 1.1: Given the multimodal pedestrian data of the user, extract the spatio-temporal features of each modality through the video encoder and the numerical encoder respectively;
[0329] Both the video encoder and the numerical encoder are single-modal encoders, and the formulas are as follows:
[0330] Video encoder:
[0331] ;
[0332] ;
[0333] Numerical encoder:
[0334] ;
[0335] ;
[0336] ;
[0337] ;
[0338] Step 1.2: Through the video encoder combined with the time-domain dependence modeling module, and the numerical encoder introducing the time-series dynamic gating mechanism, extract the spatio-temporal correlation between modalities;
[0339] The formula of the time-domain dependence modeling module is as follows:
[0340] ;
[0341] ;
[0342] The formula of the time-series dynamic gating mechanism is as follows:
[0343] ;
[0344] Step 1.3: Based on the extracted spatio-temporal features and spatio-temporal correlations, perform cross-modal feature fusion through the blending module and the Mamba module, generate mixed features using vector-level multiplication and vector-level addition operations, and capture spatio-temporal correlations through the efficient spatial scan 2D layer, and finally output the multi-modal fusion pedestrian features.
[0345] The formula of the blending module is as follows:
[0346] ;
[0347] ;
[0348] ;
[0349] ;
[0350] ;
[0351] ;
[0352] When the similarity with the upper-layer graphical representation or the similarity with the upper-layer numerical representation is greater than the set threshold, the feature embedding of the upper layer is reused.
[0353] The Mamba module generates hybrid features using vector-level multiplication and vector-level addition operations and captures spatio-temporal correlations through an efficient spatial scan 2D layer, and finally outputs multi-modal fusion pedestrian features. The formula is as follows:
[0354] ;
[0355] ;
[0356] ;
[0357] ;
[0358] ;
[0359] Further, Step 2 includes the following steps:
[0360] Step 2.1: Calculate the grid trajectory density based on the multi-modal fusion pedestrian flow features. The formula is as follows:
[0361] ;
[0362] Dynamically adjust the grid resolution through the regional feature distribution, and ensure the refined representation of sparse and dense regions by iteratively dividing the high-density regions. The formula is as follows:
[0363] ;
[0364] Step 2.2: Dynamically allocate the privacy budget for each round of iteration based on the grid resolution and the initial privacy budget, and calculate the remaining budget. Combine the regional trajectory density weights and re-allocate the remaining privacy budget;
[0365] Dynamically allocate the privacy budget for each round of iteration based on the grid resolution and the initial privacy budget. The formula is as follows:
[0366] ;
[0367] ;
[0368] ;
[0369] Calculate the remaining privacy budget, and combine the regional trajectory density weights to re - allocate the remaining privacy budget. The formula is as follows:
[0370] Calculate the remaining privacy budget:
[0371] ;
[0372] Calculate the regional trajectory density weights:
[0373] ;
[0374] Re - allocate the remaining privacy budget:
[0375] ;
[0376] Step 2.3: Add Laplace noise to the multi - modal fusion pedestrian features of each grid cell to generate perturbed features. By adjusting and balancing the noise intensity based on grid resolution and grid trajectory feature density, ensure data availability. The formula is as follows:
[0377] ;
[0378] Furthermore, step three includes the following steps:
[0379] Step 3.1: Design a sparse self - attention branch based on squared LeakyReLU. Filter out low - correlation trajectory features by non - linear activation to screen features with high query - key matching scores, avoiding redundant calculations. At the same time, utilize the negative value retention property of LeakyReLU to alleviate the problem of neuron death. The formula is as follows:
[0380] ;
[0381] Step 3.2: Construct a standard dense self - attention branch. Use Softmax to comprehensively capture potential key features in multi - modal pedestrians, preventing over - dilution of features in sparse regions. The formula is as follows:
[0382] ;
[0383] Step 3.3: Propose a dynamic weight fusion mechanism. Adaptively adjust the dual - branch attention weights through learnable parameters to balance noise suppression in dense regions and feature enhancement in sparse regions. The formula is as follows:
[0384] ;
[0385] ;
[0386] Step 3.4: Integrate the social Transformer decoder, linearly transform and encode the neighbor trajectory feature matrix into social interaction embeddings, extract pedestrian social constraints (such as avoidance and parallel behavior) through the encoder-decoder attention layer, and generate trajectory features that conform to social rationality. The formula is as follows:
[0387] ;
[0388] Step 3.5: Design a dual monitoring mechanism, combine the regression monitoring head and the classification monitoring head (outputting pedestrian behaviors and corresponding probabilities), construct a joint loss function based on the nearest neighbor clustering center, and jointly optimize the pedestrian behavior error and modal cross-entropy loss, generate the gradient of the pre-trained model. The formula is as follows:
[0389] ;
[0390] Furthermore, step four includes the following steps:
[0391] Step 4.1: Select a subset in the cloud server, calculate the original gradient calculated by each edge node using the sample s in the t-round training, and correct the original gradient through error compensation to obtain the corrected gradient. The formula is as follows:
[0392] ;
[0393] ;
[0394] ;
[0395] Step 4.2: Calculate the average gradient of each edge node, calculate the spatio-temporal correlation based on the average gradient of the previous round, and dynamically adjust the size of Top-K. The formula is as follows:
[0396] ;
[0397] ;
[0398] ;
[0399] Step 4.3: Sparsify the average gradient of each edge node through the Top-K operator to obtain the sparse gradient, and record the gradient difference as the new error compensation. The formula is as follows:
[0400] ;
[0401] ;
[0402] ;
[0403] Step 4.4: For each component of the sparse gradient, the edge node samples two variables, and then performs gradient quantization through the quantization step size. The edge node sends the quantized gradient to the cloud server. The formula is as follows:
[0404] ;
[0405] ;
[0406] ;
[0407] Further, step five includes the following steps:
[0408] Step 5.1: After the cloud server receives the quantized gradients of all edge nodes, it uses the shared random seed of edge node m to sample the same variables, and then decodes the gradients. The formula is as follows:
[0409] ;
[0410] Step 5.2: The cloud server aggregates all the decoded gradients, adjusts the weights according to the contribution degree of each edge node, updates the global model using the dynamic learning rate, and starts the next iteration. The formula is as follows:
[0411] ;
[0412] ;
[0413] ;
[0414] Step 5.3: The cloud server distributes the global model parameters to each edge node to perform the multi-modal pedestrian behavior monitoring task. In the model prediction stage, the model outputs multiple pedestrian behaviors, and uses the behavior diversity enhancement algorithm to optimize the selection strategy to cover the multi-modal pedestrian behavior patterns. The formula is as follows:
[0415] ;
[0416] .
[0417] To protect the privacy of pedestrian users' multi-modal data and model gradients, theoretical proof shows that the differential privacy protection method based on adaptive feature density perception satisfies differential privacy.
[0418] Proof: Since the perturbed multi-modal fusion embedding of pedestrian users is added with Laplace noise, where the total privacy budget is allocated as , based on the properties of differential privacy, the differential privacy protection method based on adaptive feature density perception in this example satisfies differential privacy.
[0419] Because the model perturbation gradient is added with noise that satisfies the Gaussian mechanism, where the privacy budget is allocated as and, based on the properties of relaxed differential privacy, the sparse gradient privacy protection for adaptive spatio-temporal correlation in this example satisfies differential privacy. differential privacy.
[0420] Based on the serial combination principle of differential privacy, the highly available differential privacy method for multi-modal pedestrian behavior monitoring proposed by the present invention satisfies differential privacy, where , realizing the privacy protection of multi-modal fusion embedding of local pedestrian users and model gradients.
[0421] In this embodiment, based on real pedestrian behavior datasets (ETH-UCY, SDD), different parameters are adopted: the privacy budget for local multi-modal fusion data, the privacy budget for edge model gradients, and the training batches to evaluate the availability of the privacy protection of the multi-modal perturbed data by the present invention. The results of the comparative experiments are as Figures 2 - 5 shown. The data availability comparison diagrams between the embodiments of the present invention and the traditional pedestrian behavior monitoring privacy protection methods under different total privacy budgets of different multi-modal fusion data are as Figure 2 shown; the data availability comparison diagrams between the embodiments of the present invention and the traditional pedestrian behavior monitoring privacy protection methods under different privacy budgets for edge model gradient protection are as Figure 3 shown; the average distance error comparison diagrams between the embodiments of the present invention and the traditional pedestrian behavior monitoring privacy protection methods under different training batches are as Figure 4 shown; the final distance error comparison diagrams between the embodiments of the present invention and the traditional pedestrian behavior monitoring privacy protection methods under different training batches are as Figure 5 shown. It can be seen from Figures 2 - 5 that under different privacy budgets, the data availability of this embodiment is better than that of the traditional method, and under different training batches, the average distance error and the final distance error of this embodiment are better than those of the traditional method.
[0422] Figures 2 - 5 The specific data of
[0423] .
[0424] .
[0425] .
[0426] .
[0427] It should be noted that the present invention is not limited to the above-described embodiments. The above embodiments are merely examples, and embodiments having the same constitution in terms of technical idea and achieving the same effects within the scope of the technical solution of the present invention are all included in the technical scope of the present invention. In addition, within the scope not departing from the gist of the present invention, various modifications that can be conceived by those skilled in the art to the embodiments, and other modes constructed by combining some constituent elements in the embodiments are also included in the scope of the present invention.
Claims
1. A privacy protection method for multi-modal pedestrian behavior monitoring, characterized in that, It includes the following steps: Step 1: Based on the spatio-temporal hybrid Mamba multi-modal pedestrian data fusion method, extract the characteristics of pedestrian data in each modality, and perform multi-modal spatio-temporal feature fusion; Step 2: Based on the differential privacy protection method of adaptive feature density perception, perform Laplace noise perturbation on the fused features to protect the correlation privacy of multi-modal spatio-temporal features; Step 3: Based on the multi-modal pedestrian behavior pre-training method of dual-flow dynamic Transformer, balance sparse noise features and dense noise features to improve the pre-training accuracy of the model; Step 4: A sparse gradient privacy protection method for adaptive spatio-temporal correlation, protecting the spatio-temporal correlation privacy of the model gradient; The said Step 4 includes the following steps: Step 4.1: Select a subset in the cloud server, calculate the original gradient calculated by each edge node using sample s in the t-th round of training, and correct the original gradient through error compensation to obtain the corrected gradient; Step 4.2: Calculate the average gradient of each edge node, calculate the spatio-temporal correlation according to the average gradient of the previous round, and dynamically adjust the size of Top-K; Step 4.3: Sparsify the average gradient of each edge node through the Top-K operator to obtain a sparse gradient, and record the gradient difference as the new error compensation; Step 4.4: For each component of the sparse gradient, the edge node samples two variables, and then performs gradient quantization through the quantization step size. The edge node sends the quantized gradient to the cloud server; Step 5: Based on the multi-modal pedestrian behavior monitoring method with dynamic learning rate, dynamically update the global model parameters to perform pedestrian behavior monitoring; The said Step 5 includes the following steps: Step 5.1: After the cloud server receives the quantized gradients of all edge nodes, use the shared random seed of edge node m to sample the same variables, and then decode the gradients; Step 5.2: The cloud server aggregates all decoded gradients, adjusts the weights according to the contribution degree of each edge node, updates the global model using the dynamic learning rate, and starts the next iteration; Step 5.3: The cloud server distributes the global model parameters to each edge node to perform the multi-modal pedestrian behavior monitoring task. In the model prediction stage, the model outputs multiple pedestrian behaviors, and uses the behavior diversity enhancement algorithm to optimize the selection strategy to cover multi-modal pedestrian behavior patterns.
2. The privacy protection method for multi-modal pedestrian behavior monitoring according to claim 1, characterized in that, The said Step 1 includes the following steps: Step 1.1: Given the multi-modal pedestrian data of the user, extract the spatio-temporal features of each modality through the video encoder and the numerical encoder respectively; Step 1.2: Through the video encoder combined with the time-domain dependence modeling module, the numerical encoder introduces a time-series dynamic gating mechanism to extract the spatio-temporal correlation between modalities; Step 1.3: Based on the extracted spatio-temporal features and spatio-temporal correlation, perform cross-modal feature fusion through the blending module and the Mamba module, generate mixed features using vector-level multiplication and vector-level addition operations, and capture spatio-temporal correlation through the spatial scan 2D layer, and finally output the multi-modal fusion pedestrian features.
3. A privacy protection method for multi-modal pedestrian behavior monitoring according to claim 1, characterized in that, The said Step 2 includes the following steps: Step 2.1: Calculate the grid trajectory density according to the multi-modal fusion pedestrian flow characteristics, and dynamically adjust the grid resolution through the regional feature distribution; Step 2.2: Based on the grid resolution and the initial privacy budget, dynamically allocate the privacy budget for each iteration, calculate the remaining budget, and re-allocate the remaining privacy budget in combination with the regional trajectory density weight; Step 2.3: Add Laplace noise to the multi-modal fusion pedestrian flow characteristics of each grid cell to generate perturbed characteristics, by adjusting and balancing the noise intensity based on the grid resolution and the grid trajectory feature density.
4. A privacy protection method for multi-modal pedestrian behavior monitoring according to claim 1, characterized in that, The third step includes the following steps: Step 3.1: Based on the sparse self-attention branch of squared LeakyReLU, filter out low-correlation trajectory features by non-linearly activating and screening features with high query-key matching scores; Step 3.2: Construct a standard dense self-attention branch, and use Softmax to comprehensively capture potential key features in multi-modal pedestrians; Step 3.3: Based on the dynamic weight fusion mechanism, adaptively adjust the attention weights of the two branches through learnable parameters; Step 3.4: Integrate the social Transformer decoder, linearly transform and encode the neighbor trajectory feature matrix into social interaction embeddings, extract pedestrian social constraints through the encoder-decoder attention layer, and generate trajectory features that conform to social rationality; Step 3.5: Through the dual monitoring mechanism, combine the regression monitoring head and the classification monitoring head, construct a joint loss function based on the nearest neighbor clustering center, jointly optimize the pedestrian behavior error and the modal cross-entropy loss, and generate the pre-trained model gradient.
Citation Information
Patent Citations
System and method for machine learning architecture with differential privacy
CA3097655A1
False information detection method and device for multi-modal data privacy protection
CN118965444A