Smartphone data processing method and system based on dynamic perception

CN122817841APending Publication Date: 2026-09-25深圳市阿龙电子有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611167889.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]现有技术缺点主要表现在两方面:其一,固定采样周期与静态噪声扰动策略无法适应传感信号动态变化,导致高敏感时段隐私保护不足而低敏感时段数据效用过度损失,且浅层模型对长程时序依赖的捕获能力有限,难以从复杂传感混合信号中剥离出真正反映行为状态的深层特征;其二,双分支结构的特征拼接方式缺乏动态权重调制机制,不同模态特征在融合时相互干扰,致使行为判别准确性显著下降,同时现有系统架构中各处理环节相互孤立,无法依据实时行为解析结果反向调节前端传感采集参数,形成整体处理效能的瓶颈

Benefits of technology

[0016]有益效果:本发明提出基于动态感知的智能手机数据处理方法及系统,通过移动端隐私感知扰动优化算法依据动态特征张量自身的敏感度分布实施自适应随机化扰动,解决了固定噪声策略中高敏感时段保护不足而低敏感时段数据效用过度损失的问题,有效平衡隐私保护强度与数据可用性;时序感知深度残差推断模型沿时序维度执行跨层残差递增映射并引入积分累积推断过程,增强了对长程传感信号时序依赖的捕获能力,能够从复杂混合的多源传感信号中剥离深层动态感知隐含表示,解决了浅层模型特征提取能力不足的缺陷;门控双塔时序融合函数通过门控调制机制动态权衡双塔交互卷积输出,避免不同模态特征在拼接融合时的相互干扰,提升行为状态判别的准确性与鲁棒性;手机微行为感知分析研判引擎执行多维行为显著性追踪并生成行为状态解析向量,该解析向量直接配置动态感知数据处理参数并调节后续采集策略,使前端传感采集与后端行为研判形成实时联动,解决了各处理环节相互孤立的弊端,整体上提升了智能手机动态感知处理的自适应能力、隐私保障水平及行为识别精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817841A_ABST
    Figure CN122817841A_ABST
Patent Text Reader

Abstract

The application discloses a smartphone data processing method and system based on dynamic perception, comprising: collecting multi-source sensing time sequence signals and mapping to a dynamic perception feature space to extract a multi-dimensional dynamic feature tensor, applying a random disturbance based on sensitivity analysis by a mobile terminal privacy perception disturbance optimization algorithm to generate a privacy protection dynamic feature map, inputting the dynamic feature map into a time sequence perception deep residual inference model to perform cross-layer residual incremental mapping along the time sequence dimension to extract deep dynamic perception implicit representation, using a gated double tower time sequence fusion function to perform gated modulation and double tower interactive convolution on the implicit representation to output fused time sequence dynamic features, and using a mobile phone micro-behavior perception analysis and judgment engine to perform multi-dimensional behavior saliency tracking to generate a behavior state analysis vector; the application improves the adaptive ability, privacy protection level and behavior recognition accuracy of smartphone dynamic perception processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smartphone data processing technology, and in particular to a smartphone data processing method and system based on dynamic perception. Background Technology

[0002] Modern smartphones commonly integrate various miniature sensing components such as accelerometers, gyroscopes, ambient light sensors, and electromagnetic sensors, enabling the devices to perceive user operation status and changes in the surrounding physical environment. This sensing capability provides the data foundation for intelligent human-computer interaction, behavior recognition, and context-adaptive configuration. With the continuous growth of mobile applications' demand for sensor data collection, how to extract effective features from multi-source sensor signals while protecting user privacy in dynamically changing usage scenarios has become a critical issue that urgently needs to be addressed in this technological field.

[0003] Most existing solutions employ a fixed-period sensor data sampling strategy, directly inputting the collected time-domain signals into a shallow machine learning classifier for state determination after filtering and denoising. Some solutions introduce differential privacy mechanisms, which reduce individual identifiability by superimposing Laplacian or Gaussian noise on the sensor data, and then use convolutional neural networks or long short-term memory networks to perform temporal modeling on the denoised feature sequences, finally adjusting some operating parameters of the device based on the model output. Other solutions use a dual-branch network structure to process time-domain features and spatial features separately, and then concatenate the features at the end before feeding them into a fully connected layer to complete the decision.

[0004] The shortcomings of existing technologies are mainly reflected in two aspects: First, the fixed sampling period and static noise perturbation strategy cannot adapt to the dynamic changes of sensor signals, resulting in insufficient privacy protection during highly sensitive periods and excessive loss of data utility during low-sensitivity periods. Moreover, shallow models have limited ability to capture long-range temporal dependencies, making it difficult to extract deep features that truly reflect behavioral states from complex sensor mixed signals. Second, the feature splicing method of the dual-branch structure lacks a dynamic weight modulation mechanism, and different modal features interfere with each other during fusion, resulting in a significant decrease in the accuracy of behavior discrimination. At the same time, the processing links in the existing system architecture are isolated from each other, and cannot adjust the front-end sensor acquisition parameters in reverse based on the real-time behavior analysis results, forming a bottleneck in the overall processing efficiency. Summary of the Invention

[0005] In order to overcome the shortcomings and deficiencies of existing technologies, this invention provides a smartphone data processing method and system based on dynamic perception.

[0006] The technical solution adopted in this invention is a smartphone data processing method based on dynamic perception, comprising the following steps: S1, acquiring multi-source sensor time-series signals from the smartphone, mapping the multi-source sensor time-series signals to a normalized dynamic perception feature space, and extracting a multi-dimensional dynamic feature tensor including motion mode, ambient light mode, and electromagnetic interference mode; S2, applying a randomized perturbation based on privacy sensitivity analysis to the multi-dimensional dynamic feature tensor using a mobile privacy-aware perturbation optimization algorithm to generate a privacy-protected dynamic feature map; S3, inputting the dynamic feature map into a time-series perception depth residual inference model. S4. Perform cross-layer residual incremental mapping along the temporal dimension to extract deep dynamic perception latent representation; S5. Use a gated dual-tower temporal fusion function to perform gated modulation and dual-tower interactive convolution on the deep dynamic perception latent representation to output fused temporal dynamic features; S6. Use a mobile phone micro-behavior perception analysis and judgment engine to perform multi-dimensional behavior saliency tracking on the fused temporal dynamic features to generate behavior state analysis vectors; S7. Configure the smartphone's dynamic perception data processing parameters according to the behavior state analysis vectors, and adjust the multi-source sensor temporal signal acquisition strategy for subsequent acquisition periods with the dynamic perception data processing parameters.

[0007] Furthermore, the expression for the mobile privacy-aware perturbation optimization algorithm is: ,in For the input dynamic feature tensor, For the Hadamard product operator, For a full 1 tensor, This is the perturbation step size coefficient. loss function about gradient, Let be the set of trainable weight parameters for the privacy-aware network; the privacy-preserving output of this algorithm is: ,in Let be the candidate perturbation feature tensor. The optimal perturbation output tensor. For the Frobenius norm, The covariance regularization coefficient is... For trace operation, This is for covariance matrix operations.

[0008] Furthermore, the expression for the time-aware deep residual inference model is: in This is the current time series index. For a moment The dynamic feature map input, For a moment The deep dynamic perception implicit representation. This is the length of the residual backtracking window. For backtracking offset index, For the first Learnable attention weights for each backtracking offset, It is a linear rectified activation function. For the first The linear projection weight matrix of the backtracking offset, For the first The bias vector of the backtrack offset. The cumulative residual inference process along the time-series dimension of this model is as follows: in For the duration of the points perception, For integration variables, For a moment The residual inference output features, It is a sigmoid activation function. The integral modulation weight matrix is... This is the integral modulation bias vector. .

[0009] Furthermore, the expression for the gated dual-tower timing fusion function is as follows: in This is an implicit representation for deep dynamic perception. To integrate temporal dynamic features, For the gated modulation function, the output is the gated weight tensor. This is the first pyramidal convolution transform function. This is the second pyramidal convolution transform function. For the Hadamard product operator, It is a full 1 tensor.

[0010] Furthermore, the mobile phone micro-behavior perception and analysis engine includes a field-programmable gate array (FPGA) acceleration unit and a digital signal processing (DSP) coprocessor unit. The FPGA acceleration unit is configured with a parallel pipeline architecture, and the DSP coprocessor unit is configured with a fixed-point fast Fourier transform (FFT) hard core. When the time-series-aware deep residual inference model is deployed in the FPGA acceleration unit, the residual backtracking window length is... When the gated dual-tower timing fusion function is executed in the digital signal processing coprocessor unit, the number of output channels of the gated modulation function is set to 64, and the kernel size of the dual-tower interactive convolution is set to 8. The step size is set to 2; the perturbation step size coefficient of the mobile privacy-aware perturbation optimization algorithm. At runtime, it adaptively adjusts based on the Frobenius norm of the dynamic feature tensor, with an adjustment range of 0.01 to 0.10.

[0011] Further, step S2 includes the following sub-steps: S21, performing privacy-sensitive element-wise parsing on the multidimensional dynamic feature tensor to generate sensitivity weight coefficients for each feature dimension; S22, constructing a randomized noise perturbation kernel based on the sensitivity weight coefficients, and superimposing the randomized noise perturbation kernel with the multidimensional dynamic feature tensor; S23, performing boundary constraint pruning on the superimposed tensor, and remapping elements in the pruned tensor that exceed the preset amplitude boundary to the boundary extrema; S24, outputting the pruned tensor as the privacy-protected dynamic feature map.

[0012] Further, S3 includes the following sub-steps: S31, extracting continuous time window segments from the dynamic feature map along the temporal direction to construct a temporal batch processing input queue; S32, feeding the temporal batch processing input queue layer by layer into a deep convolutional module with stacked residual connections, with each convolutional module outputting cross-layer skip connection features; S33, collecting the cross-layer skip connection features output by each deep convolutional module and performing a weighted accumulation operation to generate a cumulative residual feature map; S34, compressing the cumulative residual feature map through global temporal pooling and outputting the deep dynamic perception implicit representation.

[0013] Further, step S4 includes the following sub-steps: S41, splitting the deep dynamic perception implicit representation into a first feature branch and a second feature branch, feeding them into a gated modulation path and a dual-tower interaction path, respectively; S42, calculating the gated activation response of the first feature branch in the gated modulation path to generate a gated weight map; S43, performing dual-tower split convolution on the second feature branch in the dual-tower interaction path to extract dual-tower local temporal features and dual-tower local spatial features, respectively; S44, performing element-wise weighted integration of the gated weight map, the dual-tower local temporal features, and the dual-tower local spatial features to output the fused temporal dynamic features.

[0014] Further, step S5 includes the following sub-steps: S51, performing a sliding window scan on the fused temporal dynamic features, and extracting the peak response position and response amplitude sequence within each sliding window; S52, calculating the behavioral saliency score within each window based on the peak response position and the response amplitude sequence, and generating a saliency score trajectory; S53, performing dynamic time warping matching between the saliency score trajectory and a pre-stored behavioral template library, and outputting a matching similarity vector; S54, performing extreme value discrimination based on the matching similarity vector, and outputting the discrimination result as the behavioral state parsing vector.

[0015] A smartphone data processing system based on dynamic perception, applied to a smartphone data processing method based on dynamic perception, includes: a multi-source heterogeneous sensor signal acquisition module, used to acquire multi-source sensor time-series signals from the smartphone and map the multi-source sensor time-series signals to a normalized dynamic perception feature space, outputting a multi-dimensional dynamic feature tensor; a privacy-aware perturbation optimization encapsulator, with its input connected to the output of the multi-source heterogeneous sensor signal acquisition module, used to apply randomized perturbation to the multi-dimensional dynamic feature tensor through a mobile privacy-aware perturbation optimization algorithm, outputting a privacy-preserved dynamic feature map; and a temporal deep residual inference accelerator, with its input connected to the output of the privacy-aware perturbation optimization encapsulator, used to perform cross-layer residual incremental mapping along the temporal dimension to improve... A deep dynamic perception latent representation is extracted; a gated dual-tower dynamic fusion modulator, with its input connected to the output of the temporal deep residual inference accelerator, is used to perform gated modulation and dual-tower interactive convolution on the deep dynamic perception latent representation, and output fused temporal dynamic features; a micro-behavior semantic analysis decision engine, with its input connected to the output of the gated dual-tower dynamic fusion modulator, is used to perform multi-dimensional behavior saliency tracking on the fused temporal dynamic features, and output behavior state parsing vectors; a dynamic perception parameter adaptive configuration loop, with its input connected to the output of the micro-behavior semantic analysis decision engine and its output connected to the configuration end of the multi-source heterogeneous sensing signal acquisition module, is used to configure dynamic perception data processing parameters and adjust subsequent acquisition strategies based on the behavior state parsing vectors.

[0016] Beneficial Effects: This invention proposes a smartphone data processing method and system based on dynamic perception. By employing a mobile privacy-aware perturbation optimization algorithm that implements adaptive randomized perturbation based on the sensitivity distribution of the dynamic feature tensor itself, it solves the problem of insufficient protection during high-sensitivity periods and excessive loss of data utility during low-sensitivity periods in fixed-noise strategies, effectively balancing privacy protection strength and data availability. The time-series-aware deep residual inference model performs cross-layer residual incremental mapping along the time-series dimension and introduces an integral accumulation inference process, enhancing the ability to capture the time-series dependence of long-range sensor signals. It can extract deep dynamic perception implicit representations from complex mixed multi-source sensor signals, solving the problem of shallow model... The system addresses the deficiency in modal feature extraction capabilities. The gated dual-tower temporal fusion function dynamically balances the output of the dual-tower interactive convolution through a gated modulation mechanism, avoiding mutual interference between different modal features during splicing and fusion, thus improving the accuracy and robustness of behavioral state discrimination. The mobile phone micro-behavior perception analysis engine performs multi-dimensional behavioral saliency tracking and generates a behavioral state analytical vector. This analytical vector directly configures the dynamic perception data processing parameters and adjusts subsequent acquisition strategies, enabling real-time linkage between front-end sensing acquisition and back-end behavior judgment. This solves the drawback of isolated processing stages and overall improves the adaptive capability, privacy protection level, and behavior recognition accuracy of smartphone dynamic perception processing. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the overall process of the method of the present invention. Figure 2 This is a flowchart of method step S2 of the present invention; Figure 3 This is a flowchart of method step S3 of the present invention; Figure 4 This is a flowchart of method step S4 of the present invention; Figure 5 This is a flowchart of step S5 of the method of the present invention; Figure 6 This is a system structure diagram of the present invention. Detailed Implementation

[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] like Figure 1 As shown, a smartphone data processing method based on dynamic perception includes the following steps: S1, acquiring multi-source sensor time-series signals from the smartphone, mapping the multi-source sensor time-series signals to a normalized dynamic perception feature space, and extracting a multi-dimensional dynamic feature tensor including motion mode, ambient light mode, and electromagnetic interference mode; S2, applying a randomized perturbation based on privacy sensitivity analysis to the multi-dimensional dynamic feature tensor using a mobile terminal privacy-aware perturbation optimization algorithm to generate a privacy-protected dynamic feature map; S3, inputting the dynamic feature map into a time-series perception depth residual inference model, along the time-series dimension... S4. Perform cross-layer residual incremental mapping to extract deep dynamic perception latent representation; S5. Use a gated dual-tower temporal fusion function to perform gated modulation and dual-tower interactive convolution on the deep dynamic perception latent representation to output fused temporal dynamic features; S6. Use a mobile phone micro-behavior perception analysis and judgment engine to perform multi-dimensional behavior saliency tracking on the fused temporal dynamic features to generate behavior state analysis vectors; S7. Configure the smartphone's dynamic perception data processing parameters according to the behavior state analysis vectors, and adjust the multi-source sensor temporal signal acquisition strategy for subsequent acquisition periods with the dynamic perception data processing parameters.

[0020] During step S1, the smartphone's built-in accelerometer outputs raw triaxial acceleration signals at a sampling rate of 1600Hz, the gyroscope outputs raw triaxial angular velocity signals at a sampling rate of 800Hz, the ambient light sensor outputs raw visible light intensity signals at a sampling rate of 200Hz, and the electromagnetic interference sensor outputs raw low-frequency magnetic field fluctuation signals at a sampling rate of 400Hz. These multi-source sensor timing signals are aggregated via the device's internal high-speed bus and input into a field-programmable gate array (FPGA) buffer queue. The buffer queue depth is configured to 2048 sampling points. Subsequently, a timestamp alignment operation is performed on each sensor signal in the buffer queue, interpolating and resampling each sensor signal to a unified time reference grid according to the acquisition time. The interval of the time reference grid is set to 0.625ms, ensuring that all sensors... The signals have the same sampling density in the time dimension. The aligned multi-source sensing signals are mapped to a normalized dynamic sensing feature space, which includes a motion mode subspace, an ambient light mode subspace, and an electromagnetic interference mode subspace. The three-axis acceleration amplitude envelope and the three-axis angular velocity amplitude envelope are extracted in the motion mode subspace. The temporal fluctuation amplitude of visible light intensity is extracted in the ambient light mode subspace. The peak-to-peak value of low-frequency magnetic field fluctuation is extracted in the electromagnetic interference mode subspace. The envelopes and amplitudes extracted from the above subspaces are organized into a multi-dimensional dynamic feature tensor. The spatial dimension of this tensor is 128×128, the channel dimension is 3, corresponding to the three modes respectively, and the time dimension is 32, corresponding to 32 consecutive time windows, each with a duration of 20ms.

[0021] In step S2, the multidimensional dynamic feature tensor is input into the perturbation kernel generation module of the mobile privacy-aware perturbation optimization algorithm. This module first calculates the sensitivity resolution value for each channel of the multidimensional dynamic feature tensor. The sensitivity resolution value is calculated based on the product of the variance and kurtosis of the feature elements in each channel. The typical variance value for the motion mode channel is 0.78, and the typical kurtosis value is 2.3; the typical variance value for the ambient light mode channel is 0.45, and the typical kurtosis value is 1.9; the typical variance value for the electromagnetic interference mode channel is 0.62, and the typical kurtosis value is 3.1. Based on the sensitivity resolution value of each channel, a randomized noise perturbation kernel corresponding to each channel is constructed. The elements of the randomized noise perturbation kernel follow a Gaussian distribution with a mean of 0 and a standard deviation 0.15 times the sensitivity resolution value. The noise standard deviation of the modal channel is 0.27, the noise standard deviation of the ambient light modal channel is 0.13, and the noise standard deviation of the electromagnetic interference modal channel is 0.29. The randomized noise perturbation kernel of each channel is element-wise superimposed with the corresponding channel of the multidimensional dynamic feature tensor. The superimposed channel tensors are then subjected to boundary constraint clipping. The upper limit of the amplitude of the boundary constraint clipping is set to 2.8 times the mean of the original channel feature elements, and the lower limit is set to 0.2 times the mean of the original channel feature elements. Elements exceeding the upper limit are reassigned to the upper limit value, and elements below the lower limit are reassigned to the lower limit value. The clipped channel tensors are combined into a privacy-preserving dynamic feature map. The spatial dimension of this dynamic feature map is maintained at 128×128, the channel dimension is maintained at 3, and the time dimension is maintained at 32.

[0022] In step S3, the dynamic feature map is truncated into continuous time window segments along the time dimension. Each continuous time window segment includes 8 continuous time nodes, and the number of overlapping nodes between adjacent segments is set to 4. This constructs a temporal batch processing input queue, with a batch size of 16 segments. The temporal batch processing input queue is fed layer by layer into the temporal-aware deep residual inference model. This model includes 6 stacked deep convolutional modules with residual connections. The kernel size of the first deep convolutional module is 5×5, and the number of output channels is 64. The kernel size of the second deep convolutional module is 5×5, and the number of output channels is 128. The kernel size of the third deep convolutional module is 3×3, and the number of output channels is 128. The kernel size of the fourth deep convolutional module is 3×3, and the number of output channels is 256. The kernel size of the fifth deep convolutional module is 3×3, and the number of output channels is 256. The kernel size of the sixth deep convolutional module is... The system is 3×3 with 512 output channels. The cross-layer skip connection features output by each depth convolutional module are accumulated to subsequent modules through residual connection paths. Specifically, the skip connection features of module 2 are accumulated to module 4 after adjusting the number of channels by 1×1 convolution, and the skip connection features of module 4 are accumulated to module 6 after adjusting the number of channels by 1×1 convolution. The cross-layer skip connection features output by each depth convolutional module are collected and weighted accumulation is performed. The weighting coefficients of each module are configured as 0.05, 0.10, 0.15, 0.20, 0.25, and 0.25, respectively, to generate a cumulative residual feature map. The spatial dimension of the cumulative residual feature map is compressed to 64×64 with 512 channels. The cumulative residual feature map is then compressed by global temporal pooling with a pooling window size of 8 and a pooling stride of 4. The spatial dimension of the deep dynamic perception latent representation output after compression is 32×32 with 512 channels.

[0023] In step S4, the deep dynamic perception implicit representation is split along the channel dimension into a first feature branch and a second feature branch. Both the first and second feature branches have 256 channels. The first feature branch is fed into the gated modulation path, and the second feature branch is fed into the dual-tower interaction path. In the gated modulation path, the first feature branch sequentially passes through a global average pooling layer and two fully connected layers. The pooling window size of the global average pooling layer is 32×32. The first fully connected layer has 128 output nodes, and the second fully connected layer has 256 output nodes. The output of the second fully connected layer is mapped to the 0-1 interval using a sigmoid activation function to generate a gated weight map. The spatial dimension of the gated weight map is 32×32, and the number of channels is 256. In the dual-tower interaction path, the second feature branch is fed in parallel into a first-tower convolutional network and a second-tower convolutional network. The first convolutional network consists of two convolutional layers. The first convolutional layer has a kernel size of 3×3 and 256 output channels. The second convolutional layer has a kernel size of 3×3 and 256 output channels. The first tower convolutional network outputs dual-tower local temporal features. The second tower convolutional network consists of two convolutional layers. The first convolutional layer has a kernel size of 1×1 and 256 output channels. The second convolutional layer has a kernel size of 1×1 and 256 output channels. The second tower convolutional network outputs dual-tower local spatial features. The gated weight map is multiplied element-wise with the dual-tower local temporal features. The difference tensor between the all-1 tensor and the gated weight map is multiplied element-wise with the dual-tower local spatial features. The results of the two multiplication operations are summed element-wise to output fused temporal dynamic features. The spatial dimension of the fused temporal dynamic features is 32×32 and the number of channels is 256.

[0024] In step S5, the fused temporal dynamic features are input into the mobile phone micro-behavior perception and analysis engine. This engine first performs a sliding window scan on the fused temporal dynamic features. The width of the sliding window is set to 8 time nodes, and the sliding step size is set to 2 time nodes. Within each sliding window, the peak response position and response amplitude sequence are extracted. The peak response position is the coordinate of the element with the largest feature amplitude within that window, and the response amplitude sequence is the sequence of amplitudes of all feature elements within that window arranged in chronological order. Based on the peak response position and response amplitude sequence, the behavioral saliency score within each window is calculated. The behavioral saliency score is the product of the root mean square of the response amplitude sequence and the coordinate entropy of the peak response position. The root mean square is typically calculated as the square root of the square mean of the amplitudes within the window, and the coordinate entropy is calculated based on the frequency of the peak response position within the window, generating a saliency score trajectory. The trajectory is a time series vector with a length equal to the number of windows. The saliency score trajectory is dynamically time-warped and matched with a pre-stored behavior template library. The behavior template library includes five types of behavior templates: walking, running, stationary, going upstairs, and going downstairs. Each behavior template has a length of 64 time nodes. The curved window constraint for dynamic time warping matching is set to 12 time nodes. The output is a matching similarity vector with a dimension of 5. The value of each dimension represents the similarity between the current saliency score trajectory and the corresponding behavior template. Extreme value discrimination is performed based on the matching similarity vector. The behavior category corresponding to the dimension with the largest value in the matching similarity vector is selected as the discrimination result. The discrimination result is encoded into a behavior state parsing vector with a dimension of 5. The dimension corresponding to the discrimination result is set to 1, and the other dimensions are set to 0.

[0025] When implementing step S6, the dynamic perception data processing parameters of the smartphone are configured according to the behavior category corresponding to the dimension set to 1 in the behavior state parsing vector. If the behavior category is walking or running, the sampling rate of the accelerometer is configured to 3200Hz, the sampling rate of the gyroscope is configured to 1600Hz, the sampling rate of the ambient light sensor is configured to 100Hz, and the sampling rate of the electromagnetic interference sensor is configured to 200Hz. If the behavior category is stationary, the sampling rate of the accelerometer is configured to 800Hz, the sampling rate of the gyroscope is configured to 400Hz, the sampling rate of the ambient light sensor is configured to 400Hz, and the sampling rate of the electromagnetic interference sensor is configured to 800Hz. If the behavior category is going up or down stairs, the sampling rate of the accelerometer is configured to 2400Hz, the sampling rate of the gyroscope is configured to 1600Hz, the sampling rate of the ambient light sensor is configured to 100Hz, and the sampling rate of the electromagnetic interference sensor is configured to 200Hz. The sampling rate of the gyroscope is configured to 1200Hz, the sampling rate of the ambient light sensor is configured to 200Hz, and the sampling rate of the electromagnetic interference sensor is configured to 400Hz. The dynamic sensing data processing parameters also include the quantization bit depth of the sensor signal acquisition. The quantization bit depth is configured to 12 bits when walking or running, 16 bits when stationary, and 14 bits when going up or down stairs. The above-configured dynamic sensing data processing parameters are written into the sensor controller register of the smartphone. The sensor controller adjusts the sampling trigger frequency of each sensor and the resolution of the analog-to-digital converter in the subsequent acquisition period according to the parameter values ​​in the register. The adjusted acquisition strategy continues to act on the next acquisition cycle, and the acquisition cycle duration is set to 100ms.

[0026] Preferably, the expression for the mobile privacy-aware perturbation optimization algorithm is: ,in For the input dynamic feature tensor, For the Hadamard product operator, For a full 1 tensor, This is the perturbation step size coefficient. loss function about gradient, Let be the set of trainable weight parameters for the privacy-aware network; the privacy-preserving output of this algorithm is: ,in Let be the candidate perturbation feature tensor. The optimal perturbation output tensor. For the Frobenius norm, The covariance regularization coefficient is... For trace operation, This is for covariance matrix operations.

[0027] Specifically, the first formula of the mobile privacy-aware perturbation optimization algorithm establishes a basic perturbation mapping relationship. It uses the dynamic feature tensor as an input variable and determines the sensitive direction of perturbation application through the gradient direction of the loss function with respect to this input variable. The gradient tensor reflects the contribution of each feature element to the change in privacy loss. A perturbation step size coefficient of 0.03 achieves a balance between privacy protection strength and feature fidelity. The Hadamard multiplication operator multiplies the gradient tensor and the all-1 tensor element-wise, achieving differentiated perturbation amplitude modulation for each feature element. The all-1 tensor serves as a baseline, preserving the main information of the original features. The gradient tensor multiplied by it and then superimposed on the all-1 tensor forms the perturbation factor tensor. The Hadamard multiplication result of the original feature tensor and the perturbation factor tensor is the basic perturbation output. The second formula uses the first... The optimization framework is built based on the fundamental perturbation output of the formula. The candidate perturbation feature tensor is substituted into the first formula as an independent variable to calculate the corresponding candidate perturbation output. The squared Frobenius norm of the candidate perturbation output and the original fundamental perturbation output is used as a fidelity constraint term to ensure that the candidate perturbation does not deviate too much from the original perturbation. The introduction of the covariance regularization term is used to suppress the correlation between the dimensions of the perturbation feature tensor and avoid the perturbation from introducing systematic bias. When the covariance regularization coefficient is 0.02, it effectively reduces the redundant correlation between feature dimensions. The optimal perturbation output tensor is searched by minimizing the weighted sum of two terms. The solution process adopts gradient descent iteration to perform 300 iterations, and the step size decay rate of each iteration is 0.98, so that the output optimal perturbation tensor maintains the distribution structure of the original dynamic features and achieves differentiated privacy protection.

[0028] Preferably, the expression for the time-aware depth residual inference model is: in This is the current time series index. For a moment The dynamic feature map input, For a moment The deep dynamic perception implicit representation. This is the length of the residual backtracking window. For backtracking offset index, For the first Learnable attention weights for each backtracking offset, It is a linear rectified activation function. For the first The linear projection weight matrix of the backtracking offset, For the first The bias vector of the backtrack offset. The cumulative residual inference process along the time-series dimension of this model is as follows: in For the duration of the points perception, For integration variables, For a moment The residual inference output features, It is a sigmoid activation function. The integral modulation weight matrix is... This is the integral modulation bias vector. .

[0029] Specifically, the first formula of the time-aware deep residual inference model establishes the residual mapping relationship between the current deep features and the input features from historical moments. This is based on directly using the current dynamic feature map input as the baseline term for the residual connection, preserving the main components of the original time-series information. Simultaneously, it compensates for this by utilizing the incremental contributions generated by linear projection and nonlinear activation of the input features from each historical moment within the backtracking window. A backtracking window length of 8 covers a 160ms time-series range, sufficient to capture the periodic features of typical smartphone actions. The learnable attention weights corresponding to each backtracking offset are adaptively adjusted during model training: offset 1 has a weight of 0.12, offset 2 has a weight of 0.15, offset 3 has a weight of 0.18, offset 4 has a weight of 0.20, offset 5 has a weight of 0.15, offset 6 has a weight of 0.10, offset 7 has a weight of 0.06, and offset 8 has a weight of 0.04. The summation operation accumulates the incremental contributions from all historical moments and adds them to the baseline term. This formula uses a residual connection structure to alleviate the deep network's limitations. The vanishing gradient problem in network training enables the model to effectively utilize long-range temporal information. The second formula performs integral accumulation inference along the continuous temporal dimension based on the implicit representation output by the first formula. The integral variable traverses the continuous time interval from the current time to the integral sensing duration, which is 160ms. The implicit representation at each integration time is transformed by linear transformation and modulated by the sigmoid activation function to generate time-varying weight coefficients. These time-varying weight coefficients weight the implicit representation itself. The integral operation accumulates the weighted and modulated implicit representation over the entire time interval. This integral inference process simulates the temporal accumulation effect of biological neurons, enabling the model to capture the subsampling interval change information in the temporal signal. The discrete-time implicit representation output by the first formula is interpolated to form a continuous-time function before entering the integral operation. The derivation relationship between the two formulas is that the first formula provides the deep feature values ​​at each discrete time, and the second formula constructs these discrete values ​​into a continuous-time function and performs integral accumulation to form a complete temporal residual inference chain.

[0030] Preferably, the expression for the gated dual-tower timing fusion function is: in This is an implicit representation for deep dynamic perception. To integrate temporal dynamic features, For the gated modulation function, the output is the gated weight tensor. This is the first pyramidal convolution transform function. This is the second pyramidal convolution transform function. For the Hadamard product operator, It is a full 1 tensor.

[0031] Specifically, the gated dual-tower temporal fusion function uses the deep dynamic perception latent representation as the sole input variable. A gated modulation function outputs gate weights ranging from 0 to 1 at each spatial location of this input variable. The first and second tower convolutional functions perform feature extraction operations on this input variable with different receptive fields and kernel configurations. The first tower convolutional function uses a 3×3 kernel to extract temporally correlated features within the local neighborhood, while the second tower convolutional function uses a 1×1 kernel to extract point-by-point spatially independent features. The output features of both tower convolutions maintain the same size in both spatial and channel dimensions. The Hadamard quadrature operator multiplies the gate weight tensor with the output of the first tower convolution element-wise, highlighting the feature response region amplified by the gate weights. The full-1 tensor and the gate weight tensor... The difference tensor is multiplied element-wise with the output of the second pyramidal convolution, preserving complementary feature regions suppressed by the gate weights. The Hadamard product of the two parts is summed element-wise to form the final fused output. The core innovation of this formula is that the gate weight tensor dynamically determines the contribution ratio of the two pyramidal convolution outputs in the final fusion, rather than using a fixed splicing or average fusion strategy. The gate modulation function includes a global average pooling layer and two fully connected layers. The first fully connected layer has 128 output nodes, and the second fully connected layer has the same number of output nodes as the number of input feature channels. The sigmoid activation function maps the output of the fully connected layer to the 0 to 1 interval to form the gate weights. During training, the gate weights are adaptively adjusted according to the distribution of input features, so that the fusion strategy changes dynamically with the input, effectively solving the problem of intermodal interference.

[0032] Preferably, the mobile phone micro-behavior perception and analysis engine includes a field-programmable gate array (FPGA) acceleration unit and a digital signal processing (DSP) coprocessor unit. The FPGA acceleration unit is configured with a parallel pipeline architecture, and the DSP coprocessor unit is configured with a fixed-point fast Fourier transform (FFT) hard core. When the time-series perception deep residual inference model is deployed in the FPGA acceleration unit, the residual backtracking window length is... When the gated dual-tower timing fusion function is executed in the digital signal processing coprocessor unit, the number of output channels of the gated modulation function is set to 64, and the kernel size of the dual-tower interactive convolution is set to 8. The step size is set to 2; the perturbation step size coefficient of the mobile privacy-aware perturbation optimization algorithm. At runtime, it adaptively adjusts based on the Frobenius norm of the dynamic feature tensor, with an adjustment range of 0.01 to 0.10.

[0033] Preferred, such as Figure 2 As shown, step S2 includes the following sub-steps: S21, performing privacy-sensitive element-wise parsing on the multidimensional dynamic feature tensor to generate sensitivity weight coefficients for each feature dimension; S22, constructing a randomized noise perturbation kernel based on the sensitivity weight coefficients, and superimposing the randomized noise perturbation kernel with the multidimensional dynamic feature tensor; S23, performing boundary constraint pruning on the superimposed tensor, and remapping elements in the pruned tensor that exceed the preset amplitude boundary to the boundary extrema; S24, outputting the pruned tensor as the privacy-protected dynamic feature map.

[0034] Specifically, S21 first performs channel-by-channel privacy sensitivity analysis on the multidimensional dynamic feature tensor. For the motion modal channel, the product of the temporal variance and kurtosis of its 128 spatial locations is calculated as the raw sensitivity value. For the ambient light modal channel, the root mean square of its spatial fluctuation amplitude is calculated as the raw sensitivity value. For the electromagnetic interference modal channel, the peak value of its frequency domain power spectral density is calculated as the raw sensitivity value. The raw sensitivity values ​​of the three modes are normalized and mapped to the interval of 0.2 to 0.8 to form sensitivity weight coefficients. S22 constructs a Laplace randomized noise perturbation kernel based on the sensitivity weight coefficients of each spatial location. The location parameter of the Laplace noise is set to 0, and the scale parameter is set to 0.25 times the sensitivity weight coefficient. The noise scale value of the motion modal channel is... The noise scale ranges from 0.08 to 0.32 for the ambient light modal channel and from 0.04 to 0.16 for the electromagnetic interference modal channel. The noise scale ranges from 0.12 to 0.48 for the electromagnetic interference modal channel. The randomized noise perturbation kernel is spatially aligned with the multidimensional dynamic feature tensor and then superimposed. In step S23, the superimposed tensor is pruned by amplitude boundary constraints. The upper boundary limit is set to 3.2 times the average amplitude of the original feature at that position, and the lower boundary limit is set to 0.15 times the average amplitude of the original feature at that position. Elements exceeding the upper or lower limit are remapped to the nearest boundary extremum. In step S24, the pruned channel tensors are recombined according to the original spatial arrangement to output a privacy-preserving dynamic feature map, in which the spatial dimension and number of channels remain unchanged.

[0035] Preferred, such as Figure 3As shown, step S3 includes the following sub-steps: S31, extracting continuous time window segments from the dynamic feature map along the temporal direction to construct a temporal batch processing input queue; S32, feeding the temporal batch processing input queue layer by layer into a deep convolutional module with stacked residual connections, with each convolutional module outputting cross-layer skip connection features; S33, collecting the cross-layer skip connection features output by each deep convolutional module and performing a weighted accumulation operation to generate a cumulative residual feature map; S34, compressing the cumulative residual feature map through global temporal pooling and outputting the deep dynamic perception implicit representation.

[0036] Specifically, S31 extracts continuous time window segments from the dynamic feature map along the time dimension. The length of each segment is set to 16 consecutive time nodes, the number of overlapping nodes between adjacent segments is set to 6, and the overlap rate is 37.5%. All extracted segments are arranged in the order of acquisition to form a temporal batch processing input queue, and the batch processing size is set to 32 segments. S32 feeds this temporal batch processing input queue layer by layer into a sequence of depthwise convolutional modules stacked with residual connections. The module sequence includes 8 serial modules. The convolutional kernel size of modules 1 to 4 is 5×5, and the convolutional kernel size of modules 5 to 8 is 3×3. While each module outputs a feature map, it also transmits the feature map to the subsequent 3rd module through a skip connection path. The module's input overlay end; S33 collects the cross-layer skip connection features output by each deep convolutional module. The weighting coefficients for the 1st to 8th modules are set to 0.03, 0.06, 0.09, 0.12, 0.15, 0.18, 0.21, and 0.16, respectively. After multiplying the output features of each module by their corresponding weighting coefficients, element-wise accumulation is performed to generate a cumulative residual feature map; S34 inputs the cumulative residual feature map into the global temporal pooling layer. The pooling window size is set to 4 time nodes, and the pooling stride is set to 2 time nodes. The time dimension is compressed to 50% of the original size. After flattening, the compressed feature map outputs a deep dynamic perception latent representation with a dimension of 4096-dimensional feature vectors.

[0037] Preferred, such as Figure 4 As shown, step S4 includes the following sub-steps: S41, splitting the deep dynamic perception implicit representation into a first feature branch and a second feature branch, feeding them into a gated modulation path and a dual-tower interaction path, respectively; S42, calculating the gated activation response of the first feature branch in the gated modulation path to generate a gated weight map; S43, performing dual-tower split convolution on the second feature branch in the dual-tower interaction path to extract dual-tower local temporal features and dual-tower local spatial features, respectively; S44, performing element-wise weighted integration of the gated weight map, the dual-tower local temporal features, and the dual-tower local spatial features to output the fused temporal dynamic features.

[0038] Specifically, in S41, the deep dynamic perception latent representation is split equally along the channel dimension at a ratio of 1:1. The first feature branch and the second feature branch each obtain 2048-dimensional features. The first feature branch is guided to a gated modulation processing path, and the second feature branch is guided to a dual-tower interactive processing path. In S42, in the gated modulation path, the first feature branch compresses the 2048 dimensions to 512 dimensions through a channel compression layer, and then restores the 2048 dimensions through a channel expansion layer. The activation function between compression and expansion uses a linear rectified unit. The output of the expansion layer is mapped by a sigmoid function to generate a gated weight vector, and the values ​​of each element of this vector are distributed in the range of 0 to 1. In S43, in the dual-tower interactive path, the second feature branch is input in parallel into the first tower convolutional network. Compared to the second tower convolutional network, the first tower convolutional network consists of 3 layers of 1D convolutions with kernel lengths of 7, 5, and 3, respectively, and the number of output channels remains unchanged at 2048 dimensions. The second tower convolutional network also consists of 3 layers of 1D convolutions with kernel lengths of 3, 5, and 7, and the number of output channels remains unchanged at 2048 dimensions. The first tower network outputs a dual-tower local temporal feature vector, and the second tower network outputs a dual-tower local spatial feature vector. S44 performs element-wise multiplication of the gated weight vector with the dual-tower local temporal feature vector, performs element-wise multiplication of the difference vector between the all-1 vector and the gated weight vector with the dual-tower local spatial feature vector, and sums the two sets of multiplied vectors element-wise to output a fused temporal dynamic feature vector with a dimension of 2048.

[0039] Preferred, such as Figure 5 As shown, step S5 includes the following sub-steps: S51, performing a sliding window scan on the fused temporal dynamic features, and extracting the peak response position and response amplitude sequence within each sliding window; S52, calculating the behavioral saliency score within each window based on the peak response position and the response amplitude sequence, and generating a saliency score trajectory; S53, performing dynamic time warping matching between the saliency score trajectory and a pre-stored behavioral template library, and outputting a matching similarity vector; S54, performing extreme value discrimination based on the matching similarity vector, and outputting the discrimination result as the behavioral state parsing vector.

[0040] Specifically, S51 performs a sliding window scan along the time axis on the fused temporal dynamic feature vector. The width of the sliding window is set to 16 consecutive time sampling points, and the sliding step size is set to 4 sampling points. Within each window, the three peak response positions with the largest amplitudes and their corresponding amplitudes are extracted. The amplitudes are arranged in descending order to form a response amplitude sequence. Each window outputs a peak position coordinate triplet and an amplitude triplet. S52 calculates the position entropy value based on the peak position coordinate triplet. The position entropy value is calculated as the negative of the sum of the logarithms of the probabilities of each peak occurrence. The root mean square value of the response amplitude triplet is multiplied by the position entropy value to obtain the behavioral significance score of that window. The significance scores of all windows are concatenated in time order to generate a significance score trajectory. The trajectory length is equal to the total number of windows. S53 Input the saliency score trajectory into the dynamic time warping matcher. The pre-stored behavior template library in the matcher includes 6 behavior templates. The length of each template is 128 time sampling points. The bending path constraint range of dynamic time warping is set to 12% of the template length, i.e., 15 sampling points. The matcher outputs the inverse of the cumulative distance between the trajectory and each template as the similarity, forming a 6-dimensional matching similarity vector. S54 Perform extreme value discrimination on the matching similarity vector, select the behavior index corresponding to the dimension with the largest similarity value as the current behavior discrimination result, set the index position of the discrimination result to the value 1 in the 6-dimensional zero vector, and keep the value 0 in the other positions. This 6-dimensional one-hot encoded vector is output as the behavior state parsing vector to the subsequent configuration stage.

[0041] like Figure 6As shown, a smartphone data processing system based on dynamic perception, applied to a smartphone data processing method based on dynamic perception, includes: a multi-source heterogeneous sensor signal acquisition module, used to acquire multi-source sensor time-series signals from the smartphone and map the multi-source sensor time-series signals to a normalized dynamic perception feature space, outputting a multi-dimensional dynamic feature tensor; a privacy-aware perturbation optimization encapsulator, with its input connected to the output of the multi-source heterogeneous sensor signal acquisition module, used to apply randomized perturbation to the multi-dimensional dynamic feature tensor through a mobile privacy-aware perturbation optimization algorithm, outputting a privacy-preserved dynamic feature map; and a temporal deep residual inference accelerator, with its input connected to the output of the privacy-aware perturbation optimization encapsulator, used to perform cross-layer residual incremental mapping along the temporal dimension. The system extracts deep dynamic perception latent representations; a gated dual-tower dynamic fusion modulator, with its input connected to the output of the temporal deep residual inference accelerator, performs gated modulation and dual-tower interactive convolution on the deep dynamic perception latent representations to output fused temporal dynamic features; a micro-behavior semantic analysis decision engine, with its input connected to the output of the gated dual-tower dynamic fusion modulator, performs multi-dimensional behavior saliency tracking on the fused temporal dynamic features to output behavior state parsing vectors; and a dynamic perception parameter adaptive configuration loop, with its input connected to the output of the micro-behavior semantic analysis decision engine and its output connected to the configuration end of the multi-source heterogeneous sensing signal acquisition module, configures dynamic perception data processing parameters and adjusts subsequent acquisition strategies based on the behavior state parsing vectors.

[0042] The smartphone data processing method and system based on dynamic perception applies randomized perturbations based on sensitivity analysis to multidimensional dynamic feature tensors through a mobile privacy-aware perturbation optimization algorithm. This dynamically adjusts the perturbation intensity according to the privacy sensitivity of different feature dimensions, avoiding insufficient protection or excessive distortion under fixed noise strategies. The time-series-aware deep residual inference model performs cross-layer incremental residual mapping along the time-series dimension and introduces integral accumulation inference, enhancing the modeling ability of long-range dependencies in multi-source sensor signals. It can effectively extract deep dynamic perception implicit representations from complex mixed signals. The gated dual-tower time-series fusion function uses a gated modulation mechanism to dynamically balance the output contribution of dual-tower interactive convolution, eliminating mutual interference between different modal features during the fusion process. The mobile phone micro-behavior perception analysis and judgment engine implements multidimensional behavior saliency tracking and outputs behavior state analytical vectors. These analytical vectors directly feed back to configure dynamic perception data processing parameters to adjust subsequent acquisition strategies, enabling real-time linkage among processing stages.

[0043] To address the shortcomings of static noise disturbances in adapting to dynamic changes in sensor signals and shallow models in capturing long-range temporal dependencies, this scheme utilizes the incremental residual mapping and integral operation of the temporal-aware deep residual inference model to extract deep features along the temporal dimension. Simultaneously, a privacy-aware perturbation optimization algorithm applies differentiated perturbations based on the real-time sensitivity distribution, ensuring that the privacy protection strength matches the signal's dynamic characteristics. Furthermore, to address the deficiencies of the dual-branch structure lacking dynamic weight modulation, leading to modal interference and isolated processing stages unable to adjust acquisition parameters in reverse, this invention employs a gated modulation mechanism of a gated dual-tower temporal fusion function to weight and integrate the dual-tower outputs to suppress inter-modal interference. Simultaneously, it relies on the behavior state analysis vector to reverse-configure the front-end acquisition strategy, achieving coordinated responses in the perception, processing, and decision stages, thereby comprehensively improving adaptive capabilities and behavior recognition accuracy.

[0044] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0045] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A smartphone data processing method based on dynamic perception, characterized in that, Includes the following steps: S1, collect multi-source sensing time-series signals from a smartphone, map the multi-source sensing time-series signals to a standardized dynamic sensing feature space, and extract multi-dimensional dynamic feature tensors including motion mode, ambient light mode and electromagnetic interference mode; S2, apply a randomized perturbation based on privacy sensitivity analysis to the multidimensional dynamic feature tensor through a mobile privacy-aware perturbation optimization algorithm to generate a privacy-protected dynamic feature map; S3, input the dynamic feature map into the temporal-aware deep residual inference model, perform cross-layer residual incremental mapping along the temporal dimension, and extract the deep dynamic-aware implicit representation; S4, the deep dynamic perception implicit representation is subjected to gated modulation and dual-tower interactive convolution using a gated dual-tower temporal fusion function to output fused temporal dynamic features; S5, using the mobile phone micro-behavior perception analysis and judgment engine to perform multi-dimensional behavior saliency tracking on the fused temporal dynamic features, and generate behavior state parsing vector; S6. Configure the dynamic sensing data processing parameters of the smartphone according to the behavior state parsing vector, and adjust the multi-source sensing time sequence signal acquisition strategy for subsequent acquisition periods with the dynamic sensing data processing parameters.

2. The smartphone data processing method based on dynamic perception according to claim 1, characterized in that, The expression for the mobile privacy-aware perturbation optimization algorithm is: ,in For the input dynamic feature tensor, For the Hadamard product operator, For a full 1 tensor, This is the perturbation step size coefficient. loss function about gradient, Let be the set of trainable weight parameters for the privacy-aware network; the privacy-preserving output of this algorithm is: ,in Let be the candidate perturbation feature tensor. The optimal perturbation output tensor. For the Frobenius norm, The covariance regularization coefficient is... For trace operation, This is for covariance matrix operations.

3. The smartphone data processing method based on dynamic perception according to claim 1, characterized in that, The expression for the time-aware depth residual inference model is: in This is the current time series index. For a moment The dynamic feature map input, For a moment The deep dynamic perception implicit representation. This is the length of the residual backtracking window. For backtracking offset index, For the first Learnable attention weights for each backtracking offset, It is a linear rectified activation function. For the first The linear projection weight matrix of the backtracking offset, For the first The bias vector of the backtrack offset. The cumulative residual inference process along the time-series dimension of this model is as follows: in For the duration of the points perception, For integration variables, For a moment The residual inference output features, It is a sigmoid activation function. The integral modulation weight matrix is... This is the integral modulation bias vector. .

4. The smartphone data processing method based on dynamic perception according to claim 1, characterized in that, The expression for the gated dual-tower timing fusion function is: in This is an implicit representation for deep dynamic perception. To integrate temporal dynamic features, For the gated modulation function, the output is the gated weight tensor. This is the first pyramidal convolution transform function. This is the second pyramidal convolution transform function. For the Hadamard product operator, It is a full 1 tensor.

5. The smartphone data processing method based on dynamic perception according to claim 1, characterized in that, The mobile phone micro-behavior perception and analysis engine includes a field-programmable gate array (FPGA) acceleration unit and a digital signal processing (DSP) coprocessor. The FPGA acceleration unit is configured with a parallel pipeline architecture, and the DSP coprocessor is configured with a fixed-point fast Fourier transform (FFT) hard core. When the time-series-aware deep residual inference model is deployed in the FPGA acceleration unit, the residual backtracking window length is... When the gated dual-tower timing fusion function is executed in the digital signal processing coprocessor unit, the number of output channels of the gated modulation function is set to 64, and the kernel size of the dual-tower interactive convolution is set to 8. The step size is set to 2; the perturbation step size coefficient of the mobile privacy-aware perturbation optimization algorithm. At runtime, it adaptively adjusts based on the Frobenius norm of the dynamic feature tensor, with an adjustment range of 0.01 to 0.

10.

6. The smartphone data processing method based on dynamic perception according to claim 1, characterized in that, S2 includes: The privacy sensitivity of the multidimensional dynamic feature tensor is analyzed element by element to generate the sensitivity weight coefficients for each feature dimension. A randomized noise perturbation kernel is constructed based on the sensitivity weight coefficient. The randomized noise perturbation kernel is superimposed on the multidimensional dynamic feature tensor. Boundary constraint pruning is performed on the superimposed tensor, and elements in the pruned tensor that exceed the preset amplitude boundary are remapped to the boundary extremum. The clipped tensor is output as the privacy-preserving dynamic feature map.

7. The smartphone data processing method based on dynamic perception according to claim 1, characterized in that, S3 includes: A continuous time window segment is extracted from the dynamic feature map along the temporal direction to construct a temporal batch processing input queue; the temporal batch processing input queue is fed layer by layer into a deep convolutional module with stacked residual connections, and each convolutional module outputs cross-layer skip connection features; the cross-layer skip connection features output by each deep convolutional module are collected, and a weighted accumulation operation is performed to generate a cumulative residual feature map; The accumulated residual feature map is compressed using global temporal pooling and then output as the deep dynamic perception implicit representation.

8. The smartphone data processing method based on dynamic perception according to claim 1, characterized in that, S4 includes: The deep dynamic perception implicit representation is split into a first feature branch and a second feature branch, which are fed into the gated modulation path and the dual-tower interaction path, respectively; the gated activation response of the first feature branch is calculated in the gated modulation path to generate a gated weight map. In the dual-tower interaction path, a dual-tower separation convolution is performed on the second feature branch to extract the dual-tower local temporal features and dual-tower local spatial features respectively; the gated weight map is then integrated with the dual-tower local temporal features and dual-tower local spatial features element-wise to output the fused temporal dynamic features.

9. The smartphone data processing method based on dynamic perception according to claim 1, characterized in that, S5 includes: A sliding window scan is performed on the fused temporal dynamic features, and the peak response position and response amplitude sequence are extracted within each sliding window; the behavioral saliency score within each window is calculated based on the peak response position and the response amplitude sequence to generate a saliency score trajectory; the saliency score trajectory is dynamically time-normalized and matched with a pre-stored behavioral template library to output a matching similarity vector; S54, perform extreme value discrimination based on the matching similarity vector, and output the discrimination result as the behavior state parsing vector.

10. A smartphone data processing system based on dynamic perception, characterized in that, The system is applied to the smartphone data processing method based on dynamic perception as described in claim 1, comprising: A multi-source heterogeneous sensor signal acquisition module is used to acquire multi-source sensor time-series signals from a smartphone and map the multi-source sensor time-series signals to a normalized dynamic sensing feature space, and output a multi-dimensional dynamic feature tensor. A privacy-aware perturbation optimization encapsulator, with its input end connected to the output end of the multi-source heterogeneous sensor signal acquisition module, is used to apply randomized perturbation to the multi-dimensional dynamic feature tensor through a mobile terminal privacy-aware perturbation optimization algorithm, and output a privacy-protected dynamic feature map. The temporal deep residual inference accelerator, whose input is connected to the output of the privacy-aware perturbation optimization encapsulator, is used to perform cross-layer residual incremental mapping along the temporal dimension to extract deep dynamic perception implicit representations. A gated dual-tower dynamic fusion modulator, with its input connected to the output of the temporal deep residual inference accelerator, is used to perform gated modulation and dual-tower interactive convolution on the deep dynamic perception implicit representation and output fused temporal dynamic features. The micro-behavior semantic analysis decision engine, with its input end connected to the output end of the gated dual-tower dynamic fusion modulator, is used to perform multi-dimensional behavior saliency tracking on the fused temporal dynamic features and output a behavior state parsing vector. The dynamic sensing parameter adaptive configuration loop has its input end connected to the output end of the micro-behavior semantic analysis decision engine and its output end connected to the configuration end of the multi-source heterogeneous sensing signal acquisition module. It is used to configure the dynamic sensing data processing parameters and adjust the subsequent acquisition strategy according to the behavior state parsing vector.