Full-scene gait data adaptive acquisition and feature extraction method based on environmental parameter perception
By using environmental parameter perception methods, combined with visual and inertial sensors, and adaptively collecting and fusing gait data, the accuracy and stability issues of gait recognition in strong light environments have been solved, achieving high-precision gait recognition and anti-counterfeiting capabilities in all weather and all scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-03-31
AI Technical Summary
Existing gait recognition technologies face problems such as decreased feature extraction accuracy and insufficient recognition precision due to light interference in outdoor strong light environments. Furthermore, multimodal fusion schemes cannot adaptively adjust trust weights when switching between complex scenes, resulting in interruption of feature flow or noise interference.
By utilizing the visual sensors of mobile terminals to collect environmental optical information in real time, determining the optical micro-deformation acquisition conditions, and combining it with inertial sensor data, a weighted superposition and adaptive weighting mechanism is adopted to extract optical micro-deformation descriptors and fuse them with inertial data to generate gait feature vectors, thereby achieving gait recognition in all weather and all scenarios.
It significantly improves recognition accuracy in strong light environments, reduces system power consumption, achieves smooth transition and high-precision gait recognition in all scenarios, and has liveness detection and anti-counterfeiting capabilities.
Smart Images

Figure CN121768073A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biometrics, specifically to a method for adaptive acquisition and feature extraction of gait data across all scenarios based on environmental parameter perception. Background Technology
[0002] Currently, gait recognition, as a non-contact biometric identification technology, mainly relies on visual sensors (such as RGB cameras) or inertial sensors (such as accelerometers and gyroscopes). Visual solutions identify gait by analyzing the movement trajectory of human body contours or joints, while inertial solutions collect limb acceleration and angular velocity data through wearable devices. However, existing visual gait recognition technologies face severe challenges in outdoor high-light environments. Traditional RGB cameras are prone to localized overexposure and excessive shadow contrast under direct midday sunlight, leading to texture loss and a sharp decline in feature extraction accuracy. Simultaneously, existing active light source technologies (such as structured light and Time-of-Flight (ToF)) suffer from infrared pattern obstruction under strong outdoor sunlight, resulting in the inability to acquire effective depth or deformation information. Furthermore, active light sources consume significant power, making prolonged operation on mobile devices difficult.
[0003] On the other hand, while solutions relying solely on inertial sensors are unaffected by lighting conditions, they suffer from integral drift after prolonged operation and struggle to capture high-frequency, subtle features such as muscle tremors on the human body surface, resulting in limited discriminative power. Existing multimodal fusion solutions typically employ simple threshold switching or fixed weight splicing, lacking a deep understanding and utilization of the physical characteristics of ambient light. When transitioning from indoor shadows to outdoor bright light in complex scenarios, the system often fails to adaptively adjust the trust weights of different modalities, leading to interruptions in feature flow or the introduction of significant environmental noise, thus hindering truly seamless and continuous all-weather, all-scene recognition. Summary of the Invention
[0004] The purpose of this invention is to provide an adaptive acquisition and feature extraction method for gait data in all scenarios based on environmental parameter perception. By exploring the coherence of sunlight at a specific spatiotemporal scale, it achieves a breakthrough effect in which the recognition accuracy is significantly improved in strong light environments.
[0005] To achieve the above objectives, this invention proposes the following technical solution: a method for adaptive acquisition and feature extraction of gait data across all scenarios based on environmental parameter perception, comprising the following steps: S1. The mobile terminal's visual sensor collects environmental images in real time and calculates the spatial coherence and temporal stability of ambient light to determine whether the current environment meets the conditions for optical micro-deformation acquisition. S2, if the conditions are met, activate the high-speed continuous shooting mode to acquire the original image sequence, construct the difference pyramid and generate an enhanced interference fringe image by weighted superposition; If the conditions are not met, only inertial sensor data will be collected; S3, multi-directional filtering is performed on the enhanced interference fringe image to extract local phase information, and global phase jump constraint is introduced in combination with real-time solar elevation angle to calculate the real physical micro-deformation field on the human body surface; S4, based on micro-deformation field, generate dense point cloud, locate key points of human skeleton and extract optical micro-deformation descriptors; S5 synchronously acquires inertial sensor data and uses an adaptive weighting mechanism driven by ambient light coherence to weight and fuse the optical micro-deformation descriptor and inertial data in the transform domain to generate the final gait feature vector.
[0006] Furthermore, in this invention, in step S1, the judgment index of ambient light is calculated using a weighted fusion model, which integrates the local contrast of the image grayscale distribution, the peak decay rate of the spatial autocorrelation function, and the pixel fluctuation amplitude of the static region between consecutive frames. The peak decay rate of the spatial autocorrelation function is used to characterize the ability of incident light to form speckle interference fringes.
[0007] Furthermore, in this invention, the specific method for constructing the differential pyramid in step S2 is as follows: calculate the differential images between the current frame and previous frames at different time intervals respectively; the weighted superposition adopts an adaptive weight allocation strategy based on signal-to-noise ratio, that is, dynamically adjust the proportion of the differential images at different intervals in the linear superposition according to the signal strength of the differential images at different intervals, and then perform guided filtering on the superimposed image to smooth noise and preserve texture edges.
[0008] Furthermore, in this invention, in step S3, the extraction of local phase information adopts a multi-directional Gabor filter bank competition mechanism, and selects the direction with the largest response intensity as the main direction to extract the wrapped phase.
[0009] Furthermore, in this invention, step S3, which incorporates a global phase jump constraint based on the real-time solar altitude angle, specifically includes: calculating the current solar incidence angle using the attitude sensor and positioning information of the mobile terminal; Based on the prior theory of the maximum theoretical height difference between adjacent pixels on the human body surface, the maximum allowable phase difference threshold between adjacent pixels is calculated using the solar incidence angle. During phase unwrapping, the search range of integer periodic ambiguity is limited to a finite set of integers by using the maximum phase difference threshold, thereby eliminating the cumulative error of phase unwrapping.
[0010] Furthermore, in this invention, step S4, extracting the optical micro-deformation descriptor includes: constructing a three-dimensional spherical region of interest around each detected skeletal key point, statistically analyzing the mean normal vector distribution, covariance matrix trace, average curvature, and deformation energy relative to the initial posture of the point cloud within the region, and combining them to form a high-dimensional optical motion feature.
[0011] Furthermore, in this invention, step S5, synchronously acquiring inertial sensor data includes: acquiring triaxial acceleration, triaxial angular velocity and triaxial magnetometer data, performing downsampling and ellipsoidal fitting correction, and achieving sub-millisecond time alignment by calculating the cross-correlation peak value between the image timestamp and the inertial data timestamp.
[0012] Furthermore, in this invention, the adaptive weighting mechanism driven by ambient light coherence in step S5 specifically comprises: Construct a continuous weighting function with values ranging from zero to one. The input to this function is the ambient light spatial coherence and temporal coherence calculated in real time. The weighting function is normalized using a hyperbolic tangent nonlinear activation function. In the wavelet transform domain, the fusion ratio of the optical micro-deformation characteristic coefficient and the inertial characteristic coefficient is dynamically adjusted according to the calculated weight value. When the ambient light coherence is lower than a preset threshold, the weight of the optical feature automatically approaches zero.
[0013] Furthermore, the present invention also includes step S6: using the Teager-Kaiser energy operator to detect the start and end points of the gait cycle of the fused gait feature signal, and extracting time-domain statistical features, frequency-domain energy distribution entropy, and mutual information features of optical and inertial signals within the segmented complete gait cycle.
[0014] A mobile terminal device for implementing the above method includes: an RGB camera module, an inertial measurement unit (IMU), a memory, and a processor; the processor is configured to execute instructions for the adaptive acquisition and feature extraction method of full-scene gait data based on environmental parameter perception.
[0015] Beneficial effects: The technical solution of this application has the following technical effects: This invention proposes a method for adaptive acquisition and feature extraction of gait data across all scenarios based on environmental parameter perception. By exploiting the coherence of sunlight at specific spatiotemporal scales, it achieves a breakthrough effect by significantly improving recognition accuracy under strong light conditions. This invention eliminates the need for an additional active projection light source, utilizing the sun as a natural coherent light source. By capturing random speckle interference fringes formed on the human body surface under sunlight, it can accurately invert micro-level deformations of the skin and clothing. This mechanism not only solves the problems of overexposure and texture loss in traditional visual solutions under strong light but also transforms strong light, a source of interference, into a favorable condition for signal enhancement, greatly reducing system power consumption and achieving fully passive, high-precision measurement.
[0016] Furthermore, this invention constructs an adaptive fusion mechanism driven by ambient light coherence. By calculating the spatiotemporal coherence index of the environment in real time, it dynamically adjusts the fusion weights of optical micro-deformation features and inertial features. This design enables the system to achieve a smooth transition in various complex scenarios, from indoor low light to outdoor strong light: it automatically degenerates into stable inertial recognition under low light, and utilizes high-dimensional optical features to improve accuracy under strong light. Combined with a phase untangling constraint mechanism based on the solar altitude angle, phase ambiguity is further eliminated, ensuring the robustness of feature extraction. This invention fundamentally improves the adaptability and anti-counterfeiting security of gait recognition technology in complex natural light environments.
[0017] It should be understood that all combinations of the foregoing concepts and the additional concepts described in more detail below can be considered part of the inventive subject matter of this disclosure, provided that such concepts do not contradict each other.
[0018] The foregoing and other aspects, embodiments, and features of the teachings of the present invention will be more fully understood from the following description in conjunction with the accompanying drawings. Other additional aspects of the invention, such as features and / or beneficial effects of exemplary embodiments, will become apparent from the following description or may be learned through practice of specific embodiments according to the teachings of the present invention. Attached Figure Description
[0019] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in the various figures may be denoted by the same reference numeral. For clarity, not every component is labeled in each figure. Embodiments of various aspects of the invention will now be described by way of example and with reference to the accompanying drawings, wherein: Figure 1 This is a schematic diagram of the method flow of the present invention.
[0020] Figure 2 Comparison of speckle interference fringes on the human body surface under strong light conditions.
[0021] Figure 3 For ambient light coherenceS env ( t The curve showing the relationship between light intensity (Lux) and illumination intensity (Lux).
[0022] Figure 4 This is a gain diagram of the differential pyramid and the weighted superposition of signal-to-noise ratio (SNR). Detailed Implementation
[0023] The embodiments of the invention are described in detail below with reference to the accompanying drawings to clearly illustrate the structure, purpose, advantages, positional relationships, and connection methods of each component. It should be noted that the directional indications (such as "front," "back," "up," and "down") involved in this embodiment are based on the posture shown in the drawings and are only used to describe the relative positional relationships and movement of the components. If the posture changes, the directional indications will be adjusted accordingly. The term "connection" includes mechanical connections and electrical connections, and can be fixed connections, detachable connections, or indirect connections through an intermediate medium. The specific meaning is understood by those skilled in the art based on the context.
[0024] This embodiment is primarily based on a mobile terminal device integrating high-performance computing capabilities, such as a smartphone, augmented reality glasses, or a dedicated wearable gait analyzer. The hardware system mainly includes the following core components: Visual Acquisition Module: Employs a standard RGB CMOS camera, eliminating the need for infrared filter removal and supporting a high frame rate acquisition mode of at least 60fps. This module's function is not only to capture human form but, more importantly, to act as a "speckle interferometer detector," capturing micron-level random interference fringes formed on the human body surface under strong sunlight. Compared to traditional LiDAR or structured light modules, this module offers advantages such as zero additional power consumption, small size, and fully passive operation. Figure 2 The image shows a comparison of speckle interference fringes on the human body surface under strong light, demonstrating the core physical phenomenon: under strong sunlight, ordinary RGB cameras can indeed capture micron-level interference fringes. The left image is under indoor diffused light (no fringes), the right image is under strong direct sunlight on a sunny day (high-density random fringes), and the bottom image shows the interference fringes.
[0025] Inertial Measurement Unit (IMU): Integrates a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer, with a sampling rate set to 1000Hz. Its function is to provide basic gait data in low-light conditions such as indoors, on cloudy days, and at night, and to provide high-precision attitude reference for image acquisition.
[0026] The environmental perception and processor module includes an ambient light sensor (ALS) or a camera for light measurement, and a SoC chip integrating an NPU (Neural Processing Unit). The processor is used to run the complex AI algorithms and mathematical model solutions described in this embodiment.
[0027] Positioning and attitude calculation module: Using the absolute attitude and latitude and longitude of the GPS / BeiDou module and IMU fusion calculation device in the Earth coordinate system, it is used to calculate the solar altitude angle and azimuth angle in real time. This is the core parameter of the optical calculation of this invention.
[0028] The method flow of this embodiment mainly includes four stages: environmental perception and decision-making, optical micro-deformation acquisition, phase unwrapping and point cloud generation, and multimodal feature fusion.
[0029] Phase 1: Ambient Light Spatiotemporal Coherence Perception and Modal Decision Making The system first activates preview mode to analyze ambient light characteristics in real time to determine whether to enable the high-precision optical micro-deformation acquisition channel. To accurately distinguish between "strong direct sunlight" (with partial coherence) and "strong artificial diffused light" (without coherence), this embodiment employs an intelligent scoring model based on spatiotemporal statistical characteristics.
[0030] Processor computation time t Environmental coherence score S env ( t Its mathematical formula is as follows: ; C local Local contrast refers to the grayscale contrast in high-frequency texture areas of an image. Sunlight-generated speckle patterns have extremely high local contrast; this parameter is used for initial screening in strong light environments.
[0031] The spatial autocorrelation decay constant is obtained by calculating the spatial autocorrelation function of a local region of the image. Because the speckle particles are extremely small (micrometer scale), their autocorrelation function has an extremely narrow peak value. The autocorrelation peak is extremely small; while that of ordinary diffuse shadows is broader. Using the exponential function... It can sensitively detect whether the light field has coherence. This is the core innovation of this step, which can effectively prevent the system from being falsely triggered under indirect strong light.
[0032] Pixel temporal variance refers to the variance of pixel intensity in a static scene.
[0033] when S env ( t When the value is greater than the preset threshold T, the system determines that it is currently in a "high coherence direct sunlight environment" and activates the subsequent optical channels; otherwise, it only records inertial data to achieve optimal power consumption management.
[0034] like Figure 3 Ambient light coherence Senv ( t The relationship curve between light intensity (Lux) and illumination intensity (Lux) demonstrates the effectiveness of the intelligent triggering mechanism and quantitatively illustrates the invention. S env ( t How to distinguish between light intensity and coherence. S env ( t The rating varies with Lux; for example, 500 Lux indoors. S env ≈0, outdoor 80,000 Lux, S env ≈0.1; however, if a high-power incandescent lamp is used for illumination at 80,000 Lux, S env ≈0.1.
[0035] Phase Two: Differential Pyramid Enhancement and Interference Fringe Extraction Once the optical channel is activated, the camera captures seven frames of images continuously at 60fps. To extract faint interference fringes invisible to the naked eye from strong background light, this embodiment constructs a difference pyramid and performs adaptive weighted stacking.
[0036] Enhanced interferometric feature map E spec The calculation formula is: ; Among them, weight w k Determined dynamically by the signal-to-noise ratio (SNR): ; : Indicates the current frame and the previous frame k Frame differentiation. Differential operations can effectively remove the inherent color texture of the human body surface, i.e., the low-frequency background, and retain only the interference fringe signal that changes with micro-vibrations.
[0037] w k Adaptive weighting is used, with different deformation amplitudes captured at different time intervals k=1, 2, 3. When human movement is fast, the frame difference noise is large and the SNR is low when the interval is large, so the formula will automatically reduce its weight. w 3; and vice versa. This mechanism ensures the highest quality stripe images are obtained at different walking speeds.
[0038] like Figure 4The difference pyramid and weighted superposition signal-to-noise ratio (SNR) gain plots demonstrate the effectiveness of the signal enhancement process: quantization shows the enhancement effect of the fusion algorithm on signals with slight deformation. Three curves are compared: 1. Original single-frame difference image SNR; 2. Fixed weight superposition SNR; 3. Adaptive weighted superposition (this invention) SNR.
[0039] Linear superposition takes advantage of the fact that random noise is uncorrelated, so that the signal strength increases linearly while the noise only increases by the square root, thereby significantly improving the signal-to-noise ratio. Guided filtering further filters out sensor thermal noise while preserving the edges of the stripes.
[0040] Phase 3: Robust Phase Unwrapping Based on Solar Elevation Angle Constraints The system uses feature maps E spec Multi-directional Gabor filtering is used to extract the wrapped phase. To address the common 2π ambiguity problem in phase unwrapping, this embodiment introduces an original physical constraint mechanism.
[0041] Real-time solar zenith angle obtained using mobile phone GPS and IMU Construct phase transition constraints: ; The maximum possible physical height abrupt change between adjacent pixels on the human body surface, set based on prior knowledge of human anatomy, such as 0.5mm.
[0042] : The equivalent wavelength of the peak energy of the solar spectrum (approximately 550 nm).
[0043] Projection factor. The sensitivity of physical displacement represented by a unit phase change varies depending on the solar altitude angle.
[0044] Traditional untangling algorithms are prone to "guessing" the period number under strong noise. k This formula uses astronomical parameters to "lock" the theoretical upper limit of the phase gradient. For example, at noon when the sun is directly overhead ( The phase transition range is large, allowing for a wide range of phase transitions; however, this range narrows during oblique illumination in the early morning and late evening. This reduces the integer solution search space for phase unwrapping from infinite to {-1, 0, 1}, completely eliminating the spread of phase unwrapping errors on the human body surface, which is key to achieving sub-millimeter level accuracy.
[0045] Phase 4: Adaptive Fusion of Inertial-Optical Dual Modes After the system calculates the micro-deformation point cloud, it extracts the optical micro-deformation descriptor. Fopt and inertial characteristics F imu Integration. To achieve seamless switching around the clock, a differentiable weight adjustment mechanism based on environmental parameters was designed.
[0046] Fusion features F fused ( t The calculation is as follows: ; ; The hyperbolic tangent function, used as an activation function, maps the environmental score to the interval [0, 1). Its S-curve characteristic ensures that at critical points in lighting conditions, such as moving from shade to sunlight, the weight changes smoothly and continuously, avoiding oscillations in the recognition algorithm caused by abrupt feature changes.
[0047] S env Spatial coherence and S temp Multiplying time stability ensures that only at the optimal moment of "both bright and stable"... W ( t Only when the value approaches 1 does it dominate gait recognition; on cloudy days or indoors, W ( t The auto-recognition frequency automatically decays to 0, and the system degenerates into pure inertial recognition. This design ensures that the algorithm is robust in any environment.
[0048] The working principle of this invention is based on the speckle interference effect in physical optics. Sunlight is generally considered incoherent, but the inventors have discovered that under extremely short exposure times and specific spatial scales (direct sunlight over a clear sky), sunlight exhibits partial coherence. When this light shines on a rough human body surface (skin, fabric), microscopic interference occurs between the reflected light waves, forming a random speckle field invisible to the naked eye but perceptible to a camera. The muscle tremors and ground reaction forces during walking cause micrometer-level deformation of the surface, which directly leads to a shift in the speckle phase.
[0049] This system inverts the micro-deformation field on the human body surface by capturing this phase shift. Since the intensity of sunlight is far higher than any active light source and covers the entire body, this method is equivalent to putting a "high-sensitivity optical tactile bodysuit" on the human body, which can capture muscle tremor frequencies and skin undulation patterns that are invisible to traditional cameras, thereby extracting unique biomarkers.
[0050] This invention breaks the traditional perception that machine vision is "afraid of strong light." The stronger the light, the clearer the interference fringes, the higher the signal-to-noise ratio, and the higher the recognition accuracy. Under strong light of 80,000 Lux, the recognition rate is nearly 30% higher than traditional methods. It requires no laser emitter or structured light projection module, utilizing only existing RGB cameras and free solar energy, reducing power consumption by an order of magnitude, making it easily applicable to smartphones and smartwatches. Because the micro-deformation features originate from subcutaneous muscle tremors and bone conduction, which cannot be imitated by photos, videos, or silicone masks, it possesses extremely high liveness detection and anti-counterfeiting capabilities. Through an AI-optimized weighted formula, the system can smoothly and seamlessly switch between pure inertial mode (night / indoors) and inertial + optical enhancement mode (outdoor sunny day), solving the pain point of existing technologies being unable to simultaneously handle indoor and outdoor scenarios.
[0051] The reason why this embodiment can produce the above-mentioned technical effects is as follows: First, we can overcome the bottleneck of strong light by utilizing the "partial coherence" of sunlight.
[0052] Traditional technology considers the sun to be a broadband incoherent light source, but this only holds true on a macroscopic scale. This invention, based on the principles of physical optics, discovers that when the observation scale is reduced to the micrometer level and the exposure time is extremely short, direct sunlight reflected off rough surfaces (such as fabrics and skin) satisfies the spatial coherence condition, forming random interference patterns, i.e., speckle.
[0053] This embodiment solves the problem of information loss in traditional vision under strong light. It detects these microscopic interference fringes using the "spatial autocorrelation peak decay rate" in the formula. The stronger the sunlight, the higher the contrast of the interference fringes and the better the signal-to-noise ratio. Therefore, this invention transforms the "strong light interference" that traditional vision fears most into a "high-energy carrier wave" carrying high-frequency micro-deformation information, achieving the counterintuitive effect that the stronger the light, the better the signal quality.
[0054] Second, the mathematical gain of difference pyramids and linear superposition.
[0055] The micro-deformation signals on the human body surface are extremely weak and often drowned out by environmental background and sensor thermal noise. This embodiment solves the problem of the difficulty in extracting weak biological signals. Through calculation... The static background texture, i.e., low-frequency components, is mathematically eliminated using differential operations. The subsequent weighted linear superposition utilizes the statistical principle that "signals are coherently accumulated (amplitude increases linearly) while noise is incoherently accumulated (amplitude increases by the square root)." This significantly amplifies the interference fringe signal at specific frequencies while relatively suppressing random noise, thereby greatly improving the signal-to-noise ratio without increasing hardware costs.
[0056] Third, phase untangling constraints based on astronomical a priori knowledge.
[0057] Recovering the deformation from the interference fringes requires unwrapping the phase, which is mathematically an ill-posed problem with multiple solutions (2π ambiguity), and is particularly prone to errors in large fields of view.
[0058] This embodiment solves the global error propagation problem in phase unwrapping, i.e., the "stringing" phenomenon. This invention introduces... This physical constraint, utilizing the solar angle calculated in real-time by a mobile phone and combining it with the physical limits of skin deformation in human anatomy, allows for the calculation of the theoretical upper limit of phase change. Mathematically, this forcibly compresses the search space for phase untangling from an infinite integer domain to only three values: {-1, 0, 1}. This mathematical pruning based on physical priors completely eliminates the possibility of large deviations in the algorithm.
[0059] Fourth, ambient light-driven differentiable fusion strategy.
[0060] Existing multimodal fusion methods often involve "hard switching," which can easily lead to abrupt changes in features.
[0061] This embodiment addresses the problem of poor system stability when switching between different lighting environments. This is achieved by introducing a hyperbolic tangent function. Construct weights W ( t ). The function exhibits smooth, nonlinear saturation characteristics, mapping the physical parameters (coherence) of ambient light to continuous weights in the range [0, 1]. This means the system does not jump between "on" and "off," but rather performs continuous manifold interpolation between optical and inertial features. This ensures the differentiability of the feature vector in the time dimension, allowing the recognition model to maintain a stable gradient descent direction across all scenes and avoiding recognition oscillations caused by mode switching.
[0062] Fifth, the exploration of deep biological characteristics.
[0063] This embodiment solves the problem that traditional gait recognition is easily forged (e.g., by mimicking walking postures). This invention extracts not only the amplitude of limb swing (macroscopic movement), but also the micrometer-level tremors on the skin surface (microscopic movement) measured using optical interferometry. These tremors originate from the impact transmission of bones to the ground and the contraction of muscle fibers; they are a direct externalization of the body's internal biomechanical characteristics and are extremely difficult to imitate or forge. Therefore, this invention, while improving accuracy, also endows the system with extremely high liveness detection capabilities from a mechanistic perspective.
[0064] This embodiment also provides the following experiment: Experimental Study on Adaptive Acquisition and Feature Extraction Method of Full-Scene Gait Data Based on Environmental Parameter Awareness I. Overall Experimental Design 1.1 Experimental Objective This experiment aims to systematically verify the effectiveness of the technical solution proposed in this invention in the following aspects: (1) Improved recognition accuracy under strong light conditions; (2) The accuracy of ambient light coherence perception and modal decision-making; (3) The effectiveness of differential pyramid enhancement and phase untangling; (4) Robustness of the multimodal adaptive fusion mechanism; (5) Liveness detection and anti-counterfeiting capabilities; (6) System power consumption advantage.
[0065] 1.2 Experimental Equipment Configuration
[0066] 1.3 Construction of Experimental Dataset Data collection scale: Total number of participants: 200 (aged 18-65, male-to-female ratio 1:1) Data collection scenarios: 6 typical environments Number of walks per person per scene: 10 Total number of gait sequences: 12,000 Scene classification and lighting parameters:
[0067] II. Experiment 1: Comparison of Recognition Performance under Different Lighting Environments 2.1 Experimental Objective The accuracy of gait recognition under different lighting conditions was verified, especially the performance under strong light conditions.
[0068] 2.2 Comparison Method
[0069] 2.3 Experimental Procedure Step 1: Data Preprocessing. Each gait sequence is standardized, including: Image sequences cropped to a uniform resolution of 640×480 IMU data resampled to 1000Hz Timestamp alignment error is controlled within ±0.5ms. Step 2: Feature Extraction Gait features were extracted according to the technical approaches of each method: M1: Calculate the gait energy map (GEI) M2: Extract the trajectory of 26 key points M3: Extracts time-frequency statistics of acceleration / angular velocity M4 / M5: Extracting depth map contour features M6: Simple Feature Concatenation M7: Extract fusion features according to steps S1-S5 of this invention. Step 3: Classification and Recognition A unified recognition framework is adopted: The 200 participants were divided into a training set of 160 participants (80%) and a test set of 40 participants (20%). Five-fold cross-validation The classifiers uniformly use the ResNet-50 backbone network. Evaluation metrics: Rank-1 accuracy, equal error rate (EER) Step 4: Scenario-based testing. Test the recognition performance of each method in five scenarios, S1-S5.
[0070] 2.4 Experimental Results Table 1: Rank-1 recognition rate (%) of each method under different lighting conditions
[0071] Table 2: Equal Error Rate (EER) of each method under different lighting scenarios
[0072] 2.5 Experimental Analysis Finding 1: The present invention significantly improves performance under strong light conditions. As shown in Table 1, the method of the present invention (M7) achieves a recognition rate of 97.2% in the midday strong light scene of S5, which is about 40 percentage points higher than the traditional RGB method (M1: 52.8%, M2: 58.3%). This verifies the core technical effect of the present invention: "the stronger the light, the higher the accuracy".
[0073] Finding 2: Active light source solutions fail severely under strong light. In the S5 scene, LiDAR (M4) and structured light (M5) achieved recognition rates of only 41.2% and 35.6%, respectively, far lower than their indoor performance. This confirms the problem described in the background section where active light sources are overwhelmed by sunlight.
[0074] Discovery 3: The present invention maintains stable performance even under low light conditions. In the low-light scenario of S1, the method of this invention (79.8%) is close to that of the pure inertial method (78.4%), indicating that the adaptive fusion mechanism can automatically switch to the inertial-dominated mode in low light.
[0075] Finding 4: Fixed-weight fusion cannot adapt to changes in the scenario. The M6 method performed poorly in all scenarios, indicating that a simple fixed-weight strategy cannot cope with complex lighting changes.
[0076] III. Experiment 2: Ablation Experiment of Core Module 3.1 Experimental Objective The contribution of each core technology module of this invention is verified, including differential pyramid enhancement, solar altitude angle constraint, and adaptive fusion weight.
[0077] 3.2 Ablation Variant Design
[0078] 3.3 Experimental Results Table 3: Performance Comparison of Ablation Experiments under S5 High-Intensity Light Scenarios
[0079] Table 4: Performance Comparison of Ablation Experiments in Low-Light S1 Scene
[0080] 3.4 Experimental Analysis Analysis 1: The difference pyramid is the core source of gain. After removing the differential pyramid (A1), the strong light recognition rate plummeted from 97.2% to 76.4%, a decrease of 20.8 percentage points. This demonstrates that differential operations are crucial for extracting weak interference signals from a strong background.
[0081] Analysis 2: Solar elevation angle constraint significantly reduces phase error The A3 variant achieved a phase unwrapping error of 0.45 rad, which is 5.6 times that of the complete method (0.08 rad), and a deformation reconstruction RMSE of 12.7 μm, which is 5.5 times that of the complete method. This verifies the decisive role of physical constraints on the accuracy of phase unwrapping.
[0082] Analysis 3: Adaptive fusion weights ensure robustness in low light conditions. In the low-light scenario of S1, the recognition rate of the A5 variant (fixed weight) is only 62.4%, while the full method reaches 79.8%. This is because the fixed weight forcibly introduces low-quality optical feature noise, while the adaptive mechanism (W(t)=0.03 approaching zero) correctly suppresses the optical channel.
[0083] IV. Experiment 3: Verification of the accuracy of ambient light coherence perception 4.1 Experimental Objective Verify the discrimination accuracy of the ambient light coherence scoring model Senv(t) in step S1.
[0084] 4.2 Experimental Design Construct a test set containing 12 subdivided lighting conditions. 100 samples were collected for each condition. Manually label each sample group to determine whether it is suitable to activate the optical channel (binary classification true value). Evaluate the consistency between the model's decisions and the true values. 4.3 Subdivision of Illumination Conditions
[0085] 4.4 Experimental Results Table 5: Confusion Matrix for Environmental Coherence Judgment
[0086] Distinguishing performance indicators: Accuracy: (392+785) / 1200 = 98.1% Precision: 392 / (392 + 15) = 96.3% Recall rate: 392 / (392+8) = 98.0% F1 score: 2 × 0.963 × 0.980 / (0.963 + 0.980) = 97.1% Table 6: Discrimination accuracy under various conditions
[0087] 4.5 Experimental Analysis The model achieves slightly lower accuracy under conditions C7 (thin cloud cover) and C8 (dappled sunlight), because the spatiotemporal coherence of illumination is critical under these boundary conditions. Even so, the overall accuracy still reaches 98.1%, validating the effectiveness of the environmental perception model of this invention.
[0088] V. Experiment 4: Liveness Detection and Anti-counterfeiting Capability Verification 5.1 Experimental Objective This invention verifies its resistance to deception tactics such as photo attacks, video attacks, and imitation attacks.
[0089] 5.2 Attack Type Design
[0090] 5.3 Evaluation Indicators Attack Success Rate (ASR): The percentage of attack samples that are mistakenly identified as the target. Liveness Detection Accuracy (LDA): The accuracy rate at which a live person is correctly distinguished from an attacker. Equal Error Rate (EER): A comprehensive assessment of safety. 5.4 Experimental Results Table 7: Performance Comparison of Different Methods Against Different Attacks
[0091] 5.5 Analysis of Optical Micro-deformation Characteristics Table 8: Comparison of Micro-deformation Features between Real People and Attack Samples
[0092] 5.6 Experimental Analysis Analysis 1: Photo and video attacks are completely ineffective. This invention has a 100% resistance rate to photo and video attacks (ASR=0%). This is because planar media cannot produce true three-dimensional speckle interference fringes, and optical micro-deformation features are completely lost.
[0093] Analysis 2: Silicone and mimicry attacks were also effectively identified. Even with high-quality silicone-like clothing, the success rate of the attack is only 2.5%. This is because subcutaneous muscle tremors and bone conduction vibrations cannot be simulated by external coverings, and the spectral characteristics of micro-deformations are unique to each individual.
[0094] Analysis 3: Micro-deformation characteristics have significant distinguishing power. As shown in Table 8, the skin tremor frequency of a real person is about 8.5 Hz, while this index of various attacks is close to zero or significantly lower, forming a natural criterion for liveness detection.
[0095] VI. Experiment 5: System Power Consumption Comparison 6.1 Experimental Objective Verify the power consumption advantage of this invention compared to active light source solutions.
[0096] 6.2 Test Conditions Test duration: 30 minutes of continuous data collection Test scenario: S5 outdoor strong sunlight at noon Device battery: 4500mAh lithium battery Measurement method: Monsoon Power Monitor high-precision power consumption measurement 6.3 Experimental Results Table 9: Power Consumption Comparison of Different Methods
[0097] Table 10: Power Consumption Decomposition of Each Module in the Invention
[0098] 6.4 Experimental Analysis The power consumption of this invention is only 286mW in strong light mode, which is about 15.5% of that of LiDAR solutions and 12.2% of that of structured light solutions. This is because the invention does not require active light source emission, but only needs to use existing RGB cameras to capture the interference fringes generated by sunlight, thus fundamentally eliminating the power consumption of active illumination.
[0099] VII. Experiment 6: Stability Verification of Dynamic Scene Switching 7.1 Experimental Objective The stability of the present invention in recognizing objects during continuous switching between indoor and outdoor scenes was verified.
[0100] 7.2 Experimental Design Design a continuous walking route that incorporates various lighting conditions: Indoor office area (10m) → Glass corridor transition area (5m) → Outdoor sunshade (8m) → Outdoor direct sunlight area (15m) → Shaded area (10m) → Outdoor direct sunlight area (12m) The total length is 60m. Each participant walks 3 times, for a total of 200 people × 3 times = 600 complete tracks.
[0101] 7.3 Evaluation Indicators Sliding window recognition rate: Local recognition rate is calculated every 2 meters. Feature continuity index: cosine similarity of feature vectors of adjacent windows Weight switching smoothness: the mean of the absolute value of the first derivative of W(t) 7.4 Experimental Results Table 11: Recognition performance of each segment
[0102] Table 12: Stability Indicators at Scene Switching Boundaries
[0103] 7.5 Experimental Analysis Finding 1: The average recognition rate of this invention is significantly better than that of fixed fusion (89.8% vs. 72.0%), an improvement of 17.8 percentage points.
[0104] Finding 2: The present invention produces smoother features during scene switching with a feature jump degree of only 0.10 (0.44 for fixed fusion), indicating that the tanh smoothing mechanism of the adaptive weight W(t) effectively suppresses feature abrupt changes during mode switching.
[0105] Finding 3: Significantly reduced recognition fluctuation. The recognition fluctuation at the switching boundary of this invention is only 3.8% (compared to 20.6% for fixed fusion), which proves the effectiveness of the differentiable weighted fusion strategy.
[0106] VIII. Experiment 7: Quantitative Verification of the Constraint Effect of Solar Altitude Angle 8.1 Experimental Objective Verify the effectiveness of the phase unwrapping constraint formula under different solar altitude angles.
[0107] 8.2 Experimental Design Data was collected at different times of the day: In the morning (solar altitude angle 15°-30°) Morning (30°-60°) Noon (60°-85°) Afternoon (30°-60°) Evening (15°-30°) 100 samples are collected in each time period to evaluate the phase unwrapping error.
[0108] 8.3 Experimental Results Table 13: Phase unwrapping performance at different solar elevation angles
[0109] Table 14: Integer Period Search Space Compression Effect of Phase Untangling
[0110] 8.4 Experimental Analysis Analysis 1: The constraint mechanism is effective in all time periods. The phase unwrapping error was reduced by an average of about 80%, proving the universal effectiveness of the solar elevation angle constraint.
[0111] Analysis 2: The effect is more significant at low angles. In the morning and evening (15°-30°), the error decreases from 0.68-0.72 rad to 0.12-0.13 rad because the projection sensitivity is higher and the constraints are more stringent at this time.
[0112] Analysis 3: The search space is compressed from infinite to finite. Physical constraints compress the search space for integer periodic ambiguity from an infinite set of integers to only 2-3 candidate values, fundamentally eliminating the ambiguity of phase unwrapping.
[0113] Table 15: Comprehensive Performance Rating Table (Maximum Score: 100 points)
[0114] Based on the above seven sets of systematic experiments, the following conclusions can be drawn: Successful verification of enhanced performance under strong light: The present invention achieved a recognition rate of 97.2% in strong light environments above 80,000 Lux, which is about 40 percentage points higher than the traditional RGB method and about 60 percentage points higher than active LiDAR / structured light, verifying the core innovative effect of "the stronger the light, the higher the accuracy".
[0115] The effectiveness of the core modules has been successfully verified: ablation experiments have shown that differential pyramid enhancement contributes about 21% to the performance improvement, solar elevation angle constraint reduces phase error by about 80%, and adaptive fusion weights ensure robustness in low-light scenarios.
[0116] The accuracy of environmental perception was successfully verified: the discrimination accuracy of the ambient light coherence scoring model reached 98.1%, which can effectively distinguish between "strong direct sunlight" and "strong artificial light".
[0117] Anti-counterfeiting capabilities verified successfully: 100% resistance to photo and video attacks, 97.5% resistance to silicone attacks, and 97.8% accuracy in liveness detection.
[0118] Power consumption advantage verified: The power consumption of this invention is only 15.5% of that of active LiDAR and 12.2% of that of structured light, verifying the power consumption advantage of fully passive measurement.
[0119] Scene switching stability verification successful: In continuous indoor-outdoor switching scenarios, the feature jump degree is only 0.10 and the recognition fluctuation is only 3.8%, which is significantly better than the fixed fusion scheme.
[0120] In summary, the adaptive acquisition and feature extraction method for full-scene gait data based on environmental parameter perception proposed in this invention has been fully verified in terms of technical effectiveness. The experimental data corroborate each other, demonstrating its scientific validity and credibility.
[0121] While the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the invention. Those skilled in the art can make various modifications and refinements without departing from the spirit and scope of the invention. Therefore, the scope of protection of the present invention shall be determined by the claims.
Claims
1. A method for adaptive collection and feature extraction of full-scene gait data based on environmental parameter perception, characterized in that, The method comprises the following steps: S1, collecting environment images in real time by using a visual sensor of a mobile terminal, and calculating spatial coherence and time stability indexes of ambient light to determine whether the current environment meets optical micro-deformation collection conditions; S2, if the conditions are met, activating a high-speed continuous shooting mode to collect original image sequences, constructing a difference pyramid, and generating an enhanced interference fringe image through weighted superposition; If not, only collect inertial sensor data; S3, performing multi-directional filtering on the enhanced interference fringe image, extracting local phase information, and introducing a global phase jump constraint combined with a real-time solar elevation angle to solve a real physical micro-deformation field of a human body surface; S4, generating a dense point cloud based on the micro-deformation field, locating key points of a human body skeleton, and extracting optical micro-deformation descriptors; S5, synchronously collecting inertial sensor data, and using an adaptive weight mechanism driven by ambient light coherence to perform weighted fusion of the optical micro-deformation descriptors and the inertial data in a transform domain to generate a final gait feature vector.
2. The method of claim 1, wherein, In the step S1, the judgment index of ambient light uses a weighted fusion model, which comprehensively considers local contrast of image gray distribution, peak decay rate of spatial autocorrelation function, and pixel fluctuation amplitude of static regions between continuous frames. The peak decay rate of the spatial autocorrelation function is used to represent the ability of incident light to form speckle interference fringes.
3. The method according to claim 1 or 2, characterized in that, In the step S2, the specific method for constructing the difference pyramid is to calculate difference images between the current frame and previous frames with different time intervals respectively; the weighted superposition adopts an adaptive weight distribution strategy based on signal-to-noise ratio, that is, dynamically adjusting the proportion of different interval difference images in linear superposition according to their signal strengths, and then performing guided filtering on the superimposed image to smooth noise and retain texture edges.
4. The method according to any one of claims 1 to 3, characterized in that, In the step S3, the local phase information is extracted by using a multi-directional Gabor filter bank competition mechanism, and the direction with the strongest response intensity is selected as the main direction to extract wrapped phase.
5. The method according to any one of claims 1 to 4, characterized in that, In the step S3, the global phase jump constraint combined with the real-time solar elevation angle specifically includes: using the attitude sensor and positioning information of the mobile terminal to calculate the current solar incident angle; combined with the maximum theoretical height difference prior between adjacent pixels on the human body surface, the solar incident angle is used to calculate the maximum phase difference threshold allowed between adjacent pixels in theory; In the phase unwrapping process, the maximum phase difference threshold is used to limit the search range of integer ambiguity to a limited integer set, so as to eliminate the cumulative error of phase unwrapping.
6. The method according to any one of claims 1 to 5, characterized in that, In the step S4, the optical micro-deformation descriptor includes: constructing a three-dimensional spherical region of interest around each detected skeleton key point, and calculating the normal vector distribution mean, covariance matrix trace, average curvature, and deformation energy relative to the initial attitude of the point cloud in the region to form a high-dimensional optical motion feature.
7. The method according to any one of claims 1 to 6, characterized in that, In the step S5, synchronously collecting inertial sensor data includes: collecting three-axis acceleration, three-axis angular velocity, and three-axis magnetometer data, performing downsampling and ellipsoid fitting correction, and achieving sub-millisecond time alignment by calculating the cross-correlation peak value of the image timestamp and the inertial data timestamp.
8. The method according to any one of claims 1 to 7, characterized in that, The adaptive weight mechanism driven by ambient light coherence degree in step S5 is specifically: A continuous weight function with a value range between 0 and 1 is constructed, and the input of the function is the real-time calculated ambient light spatial coherence degree and temporal coherence degree; The weight function is normalized by using a hyperbolic tangent type nonlinear activation function; in the wavelet transform domain, the fusion ratio of the optical micro-deformation feature coefficient and the inertial feature coefficient is dynamically adjusted according to the calculated weight value, and when the ambient light coherence degree is lower than a preset threshold, the weight of the optical feature automatically tends to zero.
9. The method according to any one of claims 1 to 8, characterized in that, Further comprising step S6: using the Teager-Kaiser energy operator to detect the start point and the end point of the gait cycle for the fused gait feature signal, and extracting the time domain statistical feature, the frequency energy distribution entropy and the mutual information feature of the optical and inertial signals in the segmented complete gait cycle.
10. A mobile terminal device implementing the method of any one of claims 1 to 9, characterized in that, It comprises: An RGB camera module, an inertial measurement unit (IMU), a memory and a processor; The processor is configured to execute the instructions of the full-scene gait data adaptive acquisition and feature extraction method based on environmental parameter perception.