Privacy protection method for multi-mode pedestrian behavior monitoring

Through the multimodal data fusion of space-time hybrid Mamba, differential privacy protection of adaptive feature density perception and dual-flow dynamic Transformer's pre-training module, combined with sparse gradient privacy protection and dynamic learning rate, the problems of high computing resources, high re-identification risks and reduced data availability of pedestrian behavior data privacy protection in the existing technology are solved, and efficient and safe pedestrian behavior monitoring is achieved.

CN120012164AActive Publication Date: 2025-05-16湖南工商大学

Patent Information

Application Number
CN202510481109.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

When protecting the privacy of pedestrian behavior data, the prior art faces the problems of high computing resources, high risk of re-identification and reduced data availability, especially in high sparse pedestrian behavior data.

Method used

The multimodal pedestrian data fusion module of spatiotemporal hybrid Mamba is adopted to perform spatiotemporal multimodal pedestrian user data multimodal fusion of local multimodal pedestrian user data; the differential privacy protection module of adaptive feature density perception is used to dynamically allocate privacy budgets to protect the fusion data of multimodal pedestrian users; through the multimodal pedestrian behavior pre-training module of dual-flow dynamic Transformer, redundant and irrelevant noise characteristics are filtered, and combined with sparse gradient privacy protection and dynamic learning rate methods, highly accurate pedestrian behavior monitoring is achieved.

Benefits of technology

It effectively protects the privacy of pedestrian behavior data, improves the safety and effectiveness of multimodal pedestrian behavior monitoring, and avoids the problems of excessive computing resource consumption and reduced data availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012164A_ABST
    Figure CN120012164A_ABST
Patent Text Reader

Abstract

The invention provides a privacy protection method for multi-modal pedestrian behavior monitoring, and relates to the field of information security, and the method comprises the following steps: 1, carrying out the multi-modal pedestrian spatial-temporal feature fusion through a multi-modal pedestrian data fusion method based on spatial-temporal hybrid Mama; step 2, effectively protecting fusion features based on a differential privacy protection method of adaptive feature density perception; 3, a multi-modal pedestrian behavior pre-training method based on a double-flow dynamic Transform is adopted, and the pre-training precision of the model is improved; 4, a sparse gradient privacy protection method oriented to adaptive space-time correlation is adopted to ensure the safety of the model; and step 5, performing efficient monitoring based on a multi-modal pedestrian behavior monitoring method of a dynamic learning rate. According to the method, high-accuracy pedestrian behavior monitoring is realized, # imgabs0 # differential privacy is met based on the differential privacy property, and the safety and effectiveness of multi-mode pedestrian behavior monitoring are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information security and multimodal pedestrian behavior monitoring, and in particular to a privacy protection method for multimodal pedestrian behavior monitoring. Background Art

[0002] With the acceleration of urbanization and the popularization of smart devices, pedestrian behavior monitoring, as an important part of smart pedestrian and urban management, has gradually attracted widespread attention. Based on the cloud-edge-end collaborative architecture, pedestrian behavior monitoring technology can use distributed computing resources to efficiently process and analyze massive trajectory data, predict the movement behavior of individuals or groups in the future time period, and provide strong support for applications such as urban pedestrian flow management, public safety monitoring, and intelligent navigation. Especially in scenarios such as autonomous driving, smart travel, and emergency response, efficient and accurate pedestrian behavior monitoring can significantly improve the response speed and decision-making accuracy of the smart pedestrian system.

[0003] Although pedestrian behavior monitoring technology based on cloud-edge-end collaboration has played an important role in improving the efficiency and reliability of intelligent pedestrian systems, in reality, pedestrian behavior data is often sparse and multimodal, especially in complex urban scenarios. Pedestrian behavior data usually comes from different sensors and terminal devices, and has different collection frequencies, spatial distributions, and data quality. These data often contain sensitive personal information, such as trajectory routes, stop locations, etc. If directly uploaded to untrusted edge nodes, attackers can use link attacks, man-in-the-middle attacks, side-channel attacks, etc. to obtain personal sensitive information in user trajectories, thereby leaking user privacy.

[0004] In the prior art, researchers have proposed some privacy protection technologies, including encryption methods, anonymization technologies, differential privacy, etc., to protect the privacy of pedestrians between terminal devices and edge nodes. However, encryption methods usually require a lot of computing resources, especially when encrypting and decrypting large-scale data, which may lead to a decrease in system processing speed and response time. A core issue of anonymization technology is the risk of re-identification. Even if the data has been anonymized, attackers may still be able to recover the true identity of the anonymized data through various methods (such as association analysis, external data, etc.). The core method of differential privacy is to blur the query results by injecting noise, thereby protecting the privacy of individuals. However, existing differential privacy methods are mainly applied to intensive data. When faced with highly sparse pedestrian behavior data, they will add large noise, resulting in a significant reduction in data availability, thereby affecting the training effect and prediction accuracy of the model. In particular, in the task of pedestrian behavior monitoring, the monitoring ability of the model usually relies on a large amount of historical pedestrian behavior data and multimodal information. When these data become sparse after scrambling, the model may not be able to fully explore the regularity in the data, thereby reducing its robustness and accuracy in practical applications. Summary of the invention

[0005] The present invention is made in view of the above problems, and its purpose is to provide a privacy protection method for multimodal pedestrian behavior monitoring, a multimodal pedestrian data fusion module based on spatiotemporal hybrid Mamba, which performs spatiotemporal multimodal fusion on local multimodal pedestrian user data; a differential privacy protection module based on adaptive feature density perception, which dynamically allocates privacy budgets to protect the fusion data of multimodal pedestrian users; a multimodal pedestrian behavior pre-training module based on dual-stream dynamic Transformer, which filters redundant and irrelevant noise features and realizes multimodal pedestrian behavior monitoring pre-training; sparse gradient privacy protection for adaptive spatiotemporal correlation, which calculates the spatiotemporal correlation and dynamically adjusts the Top-K size to realize gradient sparsification and gradient spatiotemporal correlation protection; a high-availability multimodal pedestrian behavior monitoring method based on dynamic learning rate, which dynamically adjusts the weights at the learning rate and sends them to the edge nodes to realize highly accurate pedestrian behavior monitoring, and satisfies the differential privacy property. Differential privacy improves the security and effectiveness of multimodal pedestrian behavior monitoring.

[0006] Specifically, a first aspect of the present invention provides a privacy protection method for multimodal pedestrian behavior monitoring, comprising the following steps: Step 1: Based on the multimodal pedestrian data fusion method of spatiotemporal hybrid Mamba, the pedestrian data features of each modality are extracted and the multimodal spatiotemporal features are efficiently fused; Mamba is based on the Selective State Space Model (SSM) and selects input-related information by dynamically adjusting parameters. Compared with the quadratic complexity of Transformer, Mamba has linear time complexity and is particularly suitable for processing long sequence data (such as video frames and time series signals).

[0007] Multimodal pedestrian data refers to pedestrian datasets that contain multiple types of information, including videos, images, and numerical modalities.

[0008] Step 2: Based on the adaptive feature density-aware differential privacy protection method, Laplace noise perturbation is performed on the fusion features to protect the correlation privacy of multimodal spatiotemporal features; Step 3: A multimodal pedestrian behavior pre-training method based on a two-stream dynamic Transformer is used to balance sparse noise features and dense noise features to improve the model pre-training accuracy. Step 4: Sparse gradient privacy protection method for adaptive spatiotemporal correlation to protect the spatiotemporal correlation privacy of model gradients; Step 5: Based on the highly available multimodal pedestrian behavior monitoring method with dynamic learning rate, the global model parameters are dynamically updated to achieve highly accurate pedestrian behavior monitoring.

[0009] Furthermore, the step 1 comprises the following steps: Step 1.1: Given the user's multimodal pedestrian data, extract the spatiotemporal features of each modality through the image encoder and the value encoder respectively; Both the image encoder and the value encoder are single-mode encoders, and the formula is as follows: Image encoder: ; ; in: It is the image feature of the nth layer after temporal dependency modeling; is the image feature of the nth layer; is the image feature of the n-1th layer; is the layer normalization function; A module for modeling temporal dependencies; is the feed-forward layer; Numerical encoder: ; ; ; ; in: is the output of the update gate; is the sigmoid activation function; is the weight matrix of the update gate; is the hidden state of the GRU network at the previous moment; is the attention weight vector; is the Hadamard product; Is the input of the GRU network; is the output of the reset gate; is the weight matrix of the reset gate; is the weight matrix of the candidate hidden state; is the candidate hidden state; is the hyperbolic tangent function; is the hidden state of the GRU network at the current moment; Step 1.2: By combining the image encoder with the temporal dependency modeling module, the numerical encoder introduces a temporal dynamic gating mechanism to extract the spatiotemporal correlation between modalities; The time domain dependency modeling module formula is as follows: ; ; in: is the image feature of the nth layer of the gth group; is the activation function; is a learnable temporal relation parameter; is the image feature of the n-1th layer; A module for modeling temporal dependencies; is a linear transformation function; is the splicing function; is the image feature of the nth layer of the first group; The formula for the timing dynamic gating mechanism is as follows: ; in: is the attention weight vector; is the activation function; is the scoring weight matrix; is the hidden state of the GRU network at the previous moment; Is the input of the GRU network; is the dimension used to enter the value; Step 1.3: Based on the extracted spatiotemporal features and spatiotemporal correlations, cross-modal feature fusion is performed through the fusion module and the Mamba module. Hybrid features are generated using vector-level multiplication and vector-level addition operations, and spatiotemporal correlations are captured through the efficient spatial scanning 2D layer, ultimately outputting multimodal fused pedestrian features.

[0010] The formula of the fusion module is as follows: ; ; ; ; ; ; in: is the image feature embedding of the i-th layer after the self-attention layer; is the layer normalization function; is the self-attention layer; is the image view embedding of the i-1th layer; is the numerical feature embedding of the i-th layer obtained after the self-attention layer; is the numerical embedding of the i-1th layer; is the similarity with the previous layer’s image representation; It is the L2 norm, defined as the square root of the sum of the squares of the elements of the vector; is the similarity with the numerical representation of the previous layer; is the image feature embedding of layer i; is the feed-forward layer; is the i-th layer numerical feature embedding; When the similarity with the previous layer of image representation Or the similarity with the numerical representation of the previous layer When it is greater than the set threshold, the feature embedding of the previous layer is reused.

[0011] The Mamba module generates mixed features using vector-level multiplication and vector-level addition operations and captures spatiotemporal correlations through efficient spatial scanning 2D layers, and finally outputs multimodal fused pedestrian features. The formula is as follows: ; ; ; ; ; in: It is a mixed feature; It is a depth convolution operation; is the linear transformation layer; Embedding of image features; It is an element-wise multiplication operation; It is an element-wise addition operation; Embedding for numerical features; It is image-view hybrid feature embedding; is the layer normalization function; 2D layers for efficient spatial scanning; It is a numerical mixed feature embedding; Embedding for the final mixed features; To fuse pedestrian features in multi-modal manner; is the channel attention module; Furthermore, the step 2 comprises the following steps: Step 2.1: Calculate the grid trajectory density based on the multimodal fusion pedestrian flow characteristics. The formula is as follows: ; in: is the trajectory density of grid i; is the number of features of grid i; is the total number of features in the entire grid area; The grid resolution is dynamically adjusted by regional feature distribution, and the high-density area is iteratively divided to ensure the refined representation of sparse and dense areas. The formula is as follows: ; in: is the resolution vector of grid i; is the area of ​​grid i; is the trajectory density of grid i; Step 2.2: Based on the grid resolution and the initial privacy budget, dynamically allocate the privacy budget for each iteration, calculate the remaining budget, and then allocate the remaining privacy budget again based on the regional trajectory density weight. Based on the grid resolution and the initial privacy budget, the privacy budget for each iteration is dynamically allocated as follows: ; ; ; in: is the maximum number of iterations; is the regulating factor; is the partition size of the subunit; To fuse pedestrian features in multi-modal manner; To perform modulo operation on multi-modal fusion pedestrian features; is the trajectory density of grid i; is the privacy budget of the cth iteration; is the number of iterations; For predefined privacy budget; is the privacy budget of the i-th iteration; is the resolution vector of grid i; Calculate the remaining privacy budget and combine it with the regional trajectory density weight to allocate the remaining privacy budget twice. The formula is as follows: Calculate the remaining privacy budget: ; Calculate the regional trajectory density weight: ; Secondary allocation of the remaining privacy budget: ; in: Budget for remaining privacy; For predefined privacy budget; is the total number of grids; is the privacy budget of the i-th iteration; is the regional trajectory density weight of grid i; is the resolution vector of grid i; is the trajectory density of grid i; Privacy budget for quadratic distribution; Step 2.3: Add Laplace noise to the multimodal fused pedestrian features of each grid cell to generate perturbed features. Ensure data availability by adjusting and balancing the noise intensity based on grid resolution and grid trajectory feature density. The formula is as follows: ; in: is the perturbed feature of grid i; is the multimodal fusion pedestrian features of grid i; Grid-based noise added for balancing; is the Laplace noise; Privacy budget for quadratic distribution; Noise based on regional characteristics added for balancing; is the privacy budget of the i-th iteration; Furthermore, the step three comprises the following steps: Step 3.1: Design a sparse self-attention branch based on squared LeakyReLU, filter out high query-key matching score features through nonlinear activation, filter out low-correlation trajectory features, avoid redundant calculations, and use the negative value retention feature of LeakyReLU to alleviate the problem of neuron death. The formula is as follows: ; in: is a sparse self-attention branch based on squared LeakyReLU; is the activation function; is the query matrix; is the transpose of the key matrix; is the feature dimension; Step 3.2: Construct a standard dense self-attention branch and use Softmax to fully capture the potential key features of multimodal pedestrians to prevent excessive dilution of sparse area features. The formula is as follows: ; in: is the standard dense self-attention branch; is the activation function; is the query matrix; is the transpose of the key matrix; is the feature dimension; Step 3.3: Propose a dynamic weight fusion mechanism to adaptively adjust the dual-branch attention weights through learnable parameters to balance noise suppression in dense areas and feature enhancement in sparse areas. The formula is as follows: ; ; in: is the nth learnable parameter; is the nth learnable adjustment parameter; is to perform a modulo operation on the resolution vector of grid i; is the resolution vector of grid i; is the total number of grids; is the dual-branch attention weight; is a sparse self-attention branch based on squared LeakyReLU; is the standard dense self-attention branch; is the first learnable parameter; is the second learnable parameter; is the value matrix; Step 3.4: Integrate the social Transformer decoder, encode the neighbor trajectory feature matrix into social interaction embedding through linear transformation, extract pedestrian social constraints (such as avoidance and parallel behavior) through the encoder-decoder attention layer, and generate trajectory features that conform to social rationality. The formula is as follows: ; in: is the input embedding for the decoder; is a linear embedding; Embedding social constraints for pedestrians; is the learnable parameter matrix; Step 3.5: Design a dual monitoring mechanism, combining the regression monitoring head and the classification monitoring head (outputting pedestrian behaviors and corresponding probabilities), constructing a joint loss function based on the nearest neighbor cluster center, and jointly optimizing pedestrian behaviors The error and modal cross entropy loss generate the pre-trained model gradient. The formula is as follows: ; in: Gradient for the pre-trained model; is the first parameter of the balanced loss function; for Loss function; For real pedestrian behavior; To predict pedestrian behavior; is the second parameter of the balanced loss function; is the cross entropy loss function; For real pedestrian behavior The corresponding probability; To predict pedestrian behavior The corresponding probability; Furthermore, the step 4 comprises the following steps: Step 4.1: Select a subset in the cloud server, calculate the original gradient calculated by each edge node using sample s in round t training, correct the original gradient through error compensation, and obtain the corrected gradient. The formula is as follows: ; ; ; in: The original gradient calculated for edge node m using sample s in round t of training; To derive the loss function; is the model parameter of edge node m in round t; is the sample size; is the error compensation of edge node m in the tth round of training; is the average gradient of edge node m in the t-1 round of training; is the original gradient of edge node m in the t-1th round of training; is the corrected gradient; To adjust the parameters used to control the effect of error compensation on the gradient; is the maximum value function; It is the L2 norm, defined as the square root of the sum of the squares of the elements of the vector; is the clipping threshold; Step 4.2: Calculate the average gradient of each edge node, and calculate the spatiotemporal correlation based on the average gradient of the previous round, and dynamically adjust the Top-K size. The formula is as follows: ; ; ; in: is the average gradient of edge node m in round t; A small batch of data samples The number of samples; is a small batch of data samples; is the corrected gradient; is the similarity between the average gradient of the tth round and the average gradient of the t-1th round; is the average gradient of edge node m in round t-1; It is the L2 norm, defined as the square root of the sum of the squares of the elements of the vector; is the Top-K size; It is a sparse operation; To perform a modulo operation on the average gradient vector of edge node m in round t-1; Step 4.3: Use the Top-K operator to thin out the average gradient of each edge node to obtain a sparse gradient, and record the gradient difference as the new error compensation. The formula is as follows: ; ; ; in: is the average gradient of edge node m in round t after similarity calculation; is the similarity between the average gradient of the tth round and the average gradient of the t-1th round; is the average gradient of edge node m; is the sparse gradient of edge node m; It is a sparse operation; is the new error compensation of edge node m; is the gradient difference; Step 4.4: For each component of the sparse gradient, the edge node samples two variables and then quantizes the gradient by the quantization step. The edge node sends the quantized gradient to the cloud server. The formula is as follows: ; ; ; in: is the quantization step size; is the first variable sampled, ; is the component of sparse gradient; is the standard deviation of the Gaussian noise to be simulated; is the maximum value function; is a hyperparameter used to limit the lower bound of the gradient value; is the sparse gradient of edge node m in the tth round Quantity; It is the L2 norm, defined as the square root of the sum of the squares of the elements of the vector; is the quantized gradient of edge node m; is the second variable sampled, ; It is a quantization function used to map continuous gradient values ​​to discrete values; To round down; is the parameter of the function. In the second formula, x is ; Furthermore, the step five comprises the following steps: Step 5.1: After receiving the quantized gradients of all edge nodes, the cloud server uses the shared random seed of edge node m, samples the same variables, and then decodes the gradients. The formula is as follows: ; in: is the decoded quantized gradient of edge node m; is the quantized gradient of edge node m; is the second variable sampled, ; Step 5.2: The cloud server aggregates all decoded gradients, adjusts the weights according to the contribution of each edge node, updates the global model using a dynamic learning rate, and starts the next iteration. The formula is as follows: ; ; ; in: is the global gradient; is the number of edge nodes participating in the tth round of training; is the model parameter of edge node m in round t; is the gradient after decoding; is the sliding average of the squared gradient at the tth iteration; is the exponential decay rate of the squared gradient at the tth iteration; is the sliding average of the squared gradient at the t-1th iteration; is the learning rate at the tth iteration; is a constant used to prevent division by zero errors; is the global model parameter of the t+1th round; is the global model parameter of the tth round; Step 5.3: The cloud server sends the global model parameters to each edge node to perform the multimodal pedestrian behavior monitoring task. In the model prediction stage, the model outputs multiple pedestrian behaviors and uses the behavior diversity enhancement algorithm to optimize the selection strategy to cover the multimodal pedestrian behavior patterns. The formula is as follows: ; ; in: is the similarity between the behavior of pedestrian numbered i and that of pedestrian numbered j; is the time step; is the total time step; The time step is The behavior of pedestrian numbered i at time t; The time step is The behavior of the pedestrian numbered j at time t; To maximize the operation; is the total number of pedestrians; is the probability function; is the probability corresponding to the behavior of pedestrian numbered i; is the similarity between the behavior of pedestrian numbered i and that of pedestrian numbered j, and the total number of pedestrians is K; is a tuning parameter used to adjust the diversity of pedestrian behavior; is the pedestrian behavior numbered i; is the pedestrian behavior numbered j. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present drawings or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present drawings. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0013] Figure 1 is a flow chart of the steps of the present invention; Figure 2 A data availability comparison diagram of the embodiment of the present invention and the traditional pedestrian behavior monitoring privacy protection method under different total privacy budgets of multimodal fusion data; Figure 3 A data availability comparison diagram of the embodiment of the present invention and the traditional pedestrian behavior monitoring privacy protection method under different edge model gradient protection privacy budgets; Figure 4 A comparison diagram of average distance errors between the embodiment of the present invention and the traditional pedestrian behavior monitoring privacy protection method under different training batches; Figure 5 This is a comparison chart of the final distance error between the embodiment of the present invention and the traditional pedestrian behavior monitoring privacy protection method under different training batches.

[0014] The purpose, features and advantages of this figure will be further described in conjunction with the embodiments with reference to the accompanying drawings. DETAILED DESCRIPTION

[0015] In order to make the purpose, technical scheme and advantages of the present invention more clearly understood, the present invention is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments provided by the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.

[0016] Obviously, the drawings described below are only some examples or embodiments of the present invention. For ordinary technicians in this field, the present invention can also be applied to other similar scenarios based on these drawings without creative work. In addition, it can also be understood that although the efforts made in this development process may be complicated and lengthy, for ordinary technicians in this field related to the content disclosed by the present invention, some changes in design, manufacturing or production based on the technical content disclosed by the present invention are just conventional technical means, and should not be understood as insufficient content disclosed by the present invention.

[0017] If not otherwise specified, all embodiments and optional embodiments of the present invention can be combined with each other to form a new technical solution.

[0018] Unless otherwise specified, all technical features and optional technical features of the present invention can be combined with each other to form a new technical solution.

[0019] If not otherwise specified, all steps of the present invention may be performed sequentially or randomly, preferably sequentially. For example, the method includes steps (a) and (b), which means that the method may include steps (a) and (b) performed sequentially, or may include steps (b) and (a) performed sequentially. For example, the method may further include step (c), which means that step (c) may be added to the method in any order, for example, the method may include steps (a), (b) and (c), or may include steps (a), (c) and (b), or may include steps (c), (a) and (b), etc.

[0020] If there is no special explanation, the "include" and "comprising" mentioned in the present invention represent open-ended or closed-ended expressions. For example, the "include" and "comprising" may represent that other components not listed may also be included or only the listed components may be included or only the listed components may be included.

[0021] If not specifically stated, in the present invention, the term "or" is inclusive. For example, the phrase "A or B" means "A, B, or both A and B". More specifically, any of the following conditions satisfies the condition "A or B": A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); or both A and B are true (or exist).

[0022] In order to better understand the solutions of the embodiments of the present invention, some relevant terms and concepts that may be involved in the embodiments of the present invention are first introduced below.

[0023] (1) Artificial intelligence (AI), also known as intelligent machinery or machine intelligence, refers to machines that are made by humans and can display intelligence. Generally speaking, AI refers to the technology that displays human intelligence through ordinary computer programs.

[0024] (2) Deep learning (DL). Deep learning is a method of machine learning. Its concept originates from the study of artificial neural networks. Multilayer perceptrons with multiple hidden layers are a type of deep learning structure, so deep learning is often referred to as deep neural networks. Compared with general machine learning, deep learning can automatically extract features, that is, automatically combine simple features into more complex features, and use these combinations to perform multi-layer weight learning to solve problems. The motivation for studying deep learning is to establish a neural network that simulates the human brain for analysis and learning. It imitates the mechanism of the human brain to interpret data, such as images, sounds, and text. Deep learning first emerged in image recognition, but in just a few years, deep learning has been extended to various fields of machine learning and has performed well. It has been applied in various fields such as image recognition, speech recognition, audio processing, natural language recognition, robot bioinformatics processing, search engines, human-computer games, online advertising targeting, medical automatic diagnosis, and finance.

[0025] (3) A unimodal encoder is a neural network module that specifically processes a single data modality (such as text, images, and audio). Its core goal is to extract features and learn representations of input data without involving cross-modal interactions. For example, in the visual question answering (VQA) task, a unimodal encoder uses a language encoder (to process text) and a visual encoder (to process images) to extract intra-modal features through a self-attention mechanism and a feed-forward network (FFN), and then fuses them through a cross-modal encoder.

[0026] (4) Mamba is a sequence modeling architecture based on the state-space model (SSM). Its core innovations include: dynamically adjusting parameters (such as input-related matrices B, C, Δ) to decide which information to retain or ignore based on the current input, thus solving the static parameter limitations of traditional SSM; utilizing the GPU memory hierarchy to reduce memory IO overhead through scanning rather than convolution calculations, achieving linear time complexity and a 5-fold increase in inference throughput; and eliminating the need for traditional attention mechanisms, only combining SSM with multi-layer perceptron (MLP) blocks to form a simpler structure.

[0027] (5) Transformer is a deep learning model architecture for natural language processing (NLP) and other sequence-to-sequence tasks, which was first proposed by Vaswani et al. in 2017. The Transformer architecture introduces the self-attention mechanism, which is a key innovation that enables it to perform well when processing sequence data.

[0028] (6) Differential Privacy is a data protection technology that introduces random noise into the algorithm to ensure that the modification of a single data record will not significantly affect the output result, thereby protecting user privacy. It is commonly used in federated learning or sensitive data training scenarios.

[0029] In this embodiment, Figure 1 As shown, a privacy protection method for multimodal pedestrian behavior monitoring includes the following steps: Step 1: Based on the multimodal pedestrian data fusion method of spatiotemporal hybrid Mamba, the pedestrian data features of each modality are extracted and the multimodal spatiotemporal features are efficiently fused; Mamba is based on the Selective State Space Model (SSM) and selects input-related information by dynamically adjusting parameters. Compared with the quadratic complexity of Transformer, Mamba has linear time complexity and is particularly suitable for processing long sequence data (such as video frames and time series signals).

[0030] Multimodal pedestrian data refers to pedestrian datasets that contain multiple types of information, including videos, images, and numerical modalities.

[0031] Step 2: Based on the adaptive feature density-aware differential privacy protection method, Laplace noise perturbation is performed on the fusion features to protect the correlation privacy of multimodal spatiotemporal features; Step 3: A multimodal pedestrian behavior pre-training method based on a two-stream dynamic Transformer is used to balance sparse noise features and dense noise features to improve the model pre-training accuracy. Step 4: Sparse gradient privacy protection method for adaptive spatiotemporal correlation to protect the spatiotemporal correlation privacy of model gradients; Step 5: Based on the highly available multimodal pedestrian behavior monitoring method with dynamic learning rate, the global model parameters are dynamically updated to achieve highly accurate pedestrian behavior monitoring.

[0032] Furthermore, step one comprises the following steps: Step 1.1: Given the user's multimodal pedestrian data, extract the spatiotemporal features of each modality through the image encoder and the value encoder respectively; Both the image encoder and the value encoder are single-mode encoders, and the formula is as follows: Image encoder: ; ; Numerical encoder: ; ; ; ; Step 1.2: By combining the image encoder with the temporal dependency modeling module, the numerical encoder introduces a temporal dynamic gating mechanism to extract the spatiotemporal correlation between modalities; The time domain dependency modeling module formula is as follows: ; ; The formula for the timing dynamic gating mechanism is as follows: ; Step 1.3: Based on the extracted spatiotemporal features and spatiotemporal correlations, cross-modal feature fusion is performed through the fusion module and the Mamba module. Hybrid features are generated using vector-level multiplication and vector-level addition operations, and spatiotemporal correlations are captured through the efficient spatial scanning 2D layer, ultimately outputting multimodal fused pedestrian features.

[0033] The formula of the fusion module is as follows: ; ; ; ; ; ; When the similarity with the previous layer of image representation Or the similarity with the numerical representation of the previous layer When it is greater than the set threshold, the feature embedding of the previous layer is reused.

[0034] The Mamba module generates mixed features using vector-level multiplication and vector-level addition operations and captures spatiotemporal correlations through efficient spatial scanning 2D layers, and finally outputs multimodal fused pedestrian features. The formula is as follows: ; ; ; ; ; Furthermore, step 2 comprises the following steps: Step 2.1: Calculate the grid trajectory density based on the multimodal fusion pedestrian flow characteristics. The formula is as follows: ; The grid resolution is dynamically adjusted by regional feature distribution, and the high-density area is iteratively divided to ensure the refined representation of sparse and dense areas. The formula is as follows: ; Step 2.2: Based on the grid resolution and the initial privacy budget, dynamically allocate the privacy budget for each iteration, calculate the remaining budget, and then allocate the remaining privacy budget again based on the regional trajectory density weight. Based on the grid resolution and the initial privacy budget, the privacy budget for each iteration is dynamically allocated as follows: ; ; ; Calculate the remaining privacy budget and combine it with the regional trajectory density weight to allocate the remaining privacy budget twice. The formula is as follows: Calculate the remaining privacy budget: ; Calculate the regional trajectory density weight: ; Secondary allocation of the remaining privacy budget: ; Step 2.3: Add Laplace noise to the multimodal fused pedestrian features of each grid cell to generate perturbed features. Ensure data availability by adjusting and balancing the noise intensity based on grid resolution and grid trajectory feature density. The formula is as follows: ; Furthermore, step three comprises the following steps: Step 3.1: Design a sparse self-attention branch based on squared LeakyReLU, filter out high query-key matching score features through nonlinear activation, filter out low-correlation trajectory features, avoid redundant calculations, and use the negative value retention feature of LeakyReLU to alleviate the problem of neuron death. The formula is as follows: ; Step 3.2: Construct a standard dense self-attention branch and use Softmax to fully capture the potential key features of multimodal pedestrians to prevent excessive dilution of sparse area features. The formula is as follows: ; Step 3.3: Propose a dynamic weight fusion mechanism to adaptively adjust the dual-branch attention weights through learnable parameters to balance noise suppression in dense areas and feature enhancement in sparse areas. The formula is as follows: ; ; Step 3.4: Integrate the social Transformer decoder, encode the neighbor trajectory feature matrix into social interaction embedding through linear transformation, extract pedestrian social constraints (such as avoidance and parallel behavior) through the encoder-decoder attention layer, and generate trajectory features that conform to social rationality. The formula is as follows: ; Step 3.5: Design a dual monitoring mechanism, combining the regression monitoring head and the classification monitoring head (outputting pedestrian behaviors and corresponding probabilities), constructing a joint loss function based on the nearest neighbor cluster center, and jointly optimizing pedestrian behaviors The error and modal cross entropy loss generate the pre-trained model gradient. The formula is as follows: ; Furthermore, step four comprises the following steps: Step 4.1: Select a subset in the cloud server, calculate the original gradient calculated by each edge node using sample s in round t training, correct the original gradient through error compensation, and obtain the corrected gradient. The formula is as follows: ; ; ; Step 4.2: Calculate the average gradient of each edge node, and calculate the spatiotemporal correlation based on the average gradient of the previous round, and dynamically adjust the Top-K size. The formula is as follows: ; ; ; Step 4.3: Use the Top-K operator to thin out the average gradient of each edge node to obtain a sparse gradient, and record the gradient difference as the new error compensation. The formula is as follows: ; ; ; Step 4.4: For each component of the sparse gradient, the edge node samples two variables and then quantizes the gradient by the quantization step. The edge node sends the quantized gradient to the cloud server. The formula is as follows: ; ; ; Furthermore, step five comprises the following steps: Step 5.1: After receiving the quantized gradients of all edge nodes, the cloud server uses the shared random seed of edge node m, samples the same variables, and then decodes the gradients. The formula is as follows: ; Step 5.2: The cloud server aggregates all decoded gradients, adjusts the weights according to the contribution of each edge node, updates the global model using a dynamic learning rate, and starts the next iteration. The formula is as follows: ; ; ; Step 5.3: The cloud server sends the global model parameters to each edge node to perform the multimodal pedestrian behavior monitoring task. In the model prediction stage, the model outputs multiple pedestrian behaviors and uses the behavior diversity enhancement algorithm to optimize the selection strategy to cover the multimodal pedestrian behavior patterns. The formula is as follows: ; .

[0035] In order to protect the privacy of pedestrian users' multimodal data and model gradients, the differential privacy protection method based on adaptive feature density perception is theoretically proved to meet Differential privacy.

[0036] Proof: Because the perturbed multimodal fusion embedding of pedestrian users is added to satisfy the Laplace noise, the total privacy budget allocated is Based on the properties of differential privacy, the differential privacy protection method based on adaptive feature density perception in this example satisfies Differential privacy.

[0037] Because the model perturbation gradient is added to satisfy the Gaussian mechanism Noise, where the privacy budget is allocated as Based on the relaxed differential privacy property, this example is aimed at sparse gradient privacy protection with adaptive spatiotemporal correlation, satisfying Differential privacy.

[0038] Based on the serial combination principle of differential privacy, the high-availability differential privacy method for multimodal pedestrian behavior monitoring proposed in this paper meets Differential privacy, where , achieving the privacy protection of local pedestrian user multimodal fusion embedding and model gradient.

[0039] In this embodiment, based on real pedestrian behavior datasets (ETH-UCY, SDD), different parameters are used: local multimodal fusion data privacy budget, edge model gradient privacy budget and training batch to evaluate the usability of the present invention for privacy protection of multimodal perturbation data. The results of the comparative experiment are shown in Figure 2. Figure 2~Figure 5 As shown in FIG. 1 , the data availability comparison diagram of the embodiment of the present invention and the traditional pedestrian behavior monitoring privacy protection method under different total privacy budgets of multimodal fusion data is shown in FIG. Figure 2 As shown in the figure; the data availability comparison diagram of the embodiment of the present invention and the traditional pedestrian behavior monitoring privacy protection method under different edge model gradient protection privacy budgets is as follows Figure 3 As shown; the average distance error comparison diagram of the embodiment of the present invention and the traditional pedestrian behavior monitoring privacy protection method under different training batches is shown as follows Figure 4 As shown in the figure; the final distance error comparison diagram of the embodiment of the present invention and the traditional pedestrian behavior monitoring privacy protection method under different training batches is shown in Figure 5 As shown. Figure 2~Figure 5 It can be seen that the data availability of this embodiment under different privacy budgets is better than that of the traditional method, and the average distance error and final distance error of this embodiment under different training batches are better than those of the traditional method.

[0040] Figure 2~Figure 5 The specific data are shown in Table 1 to Table 4: .

[0041] .

[0042] .

[0043] .

[0044] It should be noted that the present invention is not limited to the above-mentioned embodiments. The above-mentioned embodiments are only examples, and the embodiments having the same structure as the technical idea and exerting the same effect within the scope of the technical solution of the present invention are all included in the technical scope of the present invention. In addition, without departing from the scope of the main purpose of the present invention, various modifications that can be thought of by those skilled in the art to the embodiments and other methods of combining some of the constituent elements in the embodiments are also included in the scope of the present invention.

Claims

1. A privacy protection method for multimodal pedestrian behavior monitoring, characterized in that: The following steps are involved: Step 1: Based on the multimodal pedestrian data fusion method of spatiotemporal hybrid Mamba, the pedestrian data features of each modality are extracted and multimodal spatiotemporal feature fusion is performed; Step 2: Based on the adaptive feature density-aware differential privacy protection method, Laplace noise perturbation is performed on the fusion features to protect the correlation privacy of multimodal spatiotemporal features; Step 3: A multimodal pedestrian behavior pre-training method based on a two-stream dynamic Transformer is used to balance sparse noise features and dense noise features to improve the model pre-training accuracy. Step 4: Sparse gradient privacy protection method for adaptive spatiotemporal correlation to protect the spatiotemporal correlation privacy of model gradients; Step 5: Based on the multimodal pedestrian behavior monitoring method with dynamic learning rate, the global model parameters are dynamically updated to perform pedestrian behavior monitoring.

2. According to claim 1, a privacy protection method for multimodal pedestrian behavior monitoring is characterized in that: The step 1 comprises the following steps: Step 1.1: Given the user's multimodal pedestrian data, extract the spatiotemporal features of each modality through the image encoder and the value encoder respectively; Step 1.2: By combining the image encoder with the temporal dependency modeling module, the numerical encoder introduces a temporal dynamic gating mechanism to extract the spatiotemporal correlation between modalities; Step 1.3: Based on the extracted spatiotemporal features and spatiotemporal correlations, cross-modal feature fusion is performed through the fusion module and the Mamba module. The hybrid features are generated by vector-level multiplication and vector-level addition operations, and the spatiotemporal correlations are captured through the spatial scanning 2D layer. Finally, the multimodal fused pedestrian features are output.

3. The privacy protection method for multimodal pedestrian behavior monitoring according to claim 1 is characterized in that: The step 2 comprises the following steps: Step 2.1: Calculate the grid trajectory density based on the multimodal fusion pedestrian flow characteristics, and dynamically adjust the grid resolution based on the regional feature distribution; Step 2.2: Based on the grid resolution and the initial privacy budget, dynamically allocate the privacy budget for each iteration, calculate the remaining budget, and then allocate the remaining privacy budget again based on the regional trajectory density weight. Step 2.3: Add Laplace noise to the multimodal fused pedestrian flow features of each grid cell to generate perturbed features by adjusting and balancing the noise intensity based on the grid resolution and based on the grid trajectory feature density.

4. The privacy protection method for multimodal pedestrian behavior monitoring according to claim 1 is characterized in that: The step three comprises the following steps: Step 3.1: The sparse self-attention branch based on squared LeakyReLU filters out features with high query-key matching scores through nonlinear activation and removes low-relevance trajectory features; Step 3.2: Construct a standard dense self-attention branch and use Softmax to fully capture the potential key features of multimodal pedestrians; Step 3.3: Based on the dynamic weight fusion mechanism, the dual-branch attention weights are adaptively adjusted through learnable parameters; Step 3.4: Integrate the social Transformer decoder, encode the neighbor trajectory feature matrix into social interaction embedding through linear transformation, extract pedestrian social constraints through the encoder-decoder attention layer, and generate trajectory features that conform to social rationality; Step 3.5: Through the dual monitoring mechanism, combining the regression monitoring head and the classification monitoring head, construct a joint loss function based on the nearest neighbor cluster center, jointly optimize the pedestrian behavior error and the modal cross entropy loss, and generate the pre-trained model gradient.

5. The privacy protection method for multimodal pedestrian behavior monitoring according to claim 1 is characterized in that: The step 4 comprises the following steps: Step 4.1: Select a subset in the cloud server, calculate the original gradient calculated by each edge node using sample s in round t training, correct the original gradient through error compensation, and obtain the corrected gradient; Step 4.2: Calculate the average gradient of each edge node, and calculate the spatiotemporal correlation based on the average gradient of the previous round, and dynamically adjust the Top-K size; Step 4.3: Use the Top-K operator to thin out the average gradient of each edge node to obtain a sparse gradient, and record the gradient difference as a new error compensation; Step 4.4: For each component of the sparse gradient, the edge node samples two variables and then quantizes the gradient by the quantization step size. The edge node sends the quantized gradient to the cloud server.

6. The privacy protection method for multimodal pedestrian behavior monitoring according to claim 1 is characterized in that: The step five comprises the following steps: Step 5.1: After receiving the quantized gradients of all edge nodes, the cloud server uses the shared random seed of edge node m to sample the same variables and then decode the gradients; Step 5.2: The cloud server aggregates all decoded gradients, adjusts the weights according to the contribution of each edge node, updates the global model using a dynamic learning rate, and starts the next iteration; Step 5.3: The cloud server sends the global model parameters to each edge node to perform the multimodal pedestrian behavior monitoring task. In the model prediction stage, the model outputs multiple pedestrian behaviors and uses the behavior diversity enhancement algorithm to optimize the selection strategy to cover the multimodal pedestrian behavior patterns.

Citation Information

Patent Citations

  • System and method for machine learning architecture with differential privacy

    CA3097655A1

  • Social activity recommendation method and device for multi-modal data privacy protection

    CN117540106A

  • False information detection method and device for multi-modal data privacy protection

    CN118965444A

  • Intelligent security and protection monitoring method and system based on image recognition

    CN119495054A

  • Community discovery method and device for multi-modal spatio-temporal correlation privacy protection

    CN119760451A

Cited By

  • Video stream data-oriented emotion recognition privacy protection method and system

    CN120257369A

  • Network space surveying and mapping threat detection method and system based on causal association privacy protection

    CN120337301A

  • Cloud edge-end collaborative scheduling optimization method and system based on dynamic resource portrait

    CN120803628A

  • A cloud-edge-end collaborative scheduling optimization method and system based on dynamic resource profiling

    CN120803628B

  • Multi-modal attack behavior adaptive identification privacy protection method and system

    CN121959636A