Long-distance identity recognition method for large area
Through gait recognition and multi-target cross-camera tracking technology, the problems of identity recognition and trajectory tracking in large areas are solved, efficient and accurate identity recognition and security management are achieved, and real-time monitoring and early warning capabilities are provided in complex environments.
Patent Information
- Application Number
- CN202510746806.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional identity recognition technology lacks cross-camera tracking and information fusion in large industrial parks and public places, making it difficult to achieve accurate personnel identification and trajectory tracking, especially in complex environments.
It adopts the individual's unique walking posture (gait recognition) combined with multi-target cross-camera tracking technology, uses deep learning and the gait multi-task recognition model (MLDT) for identity recognition, and realizes identity recognition without the active cooperation of personnel through gait feature extraction and multi-camera trajectory fusion.
It achieves efficient and accurate long-distance identity recognition and personnel trajectory tracking in complex environments, improves the security management level of large areas, and provides real-time monitoring and early warning capabilities for illegal intrusions.
Smart Images

Figure CN120656207A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a long-distance identity recognition method for a large area. Background Art
[0002] With the continuous development of society and the economy, large-scale industrial parks (such as chemical parks and biopharmaceutical parks) and large public places (such as airports, train stations, and large event venues) continue to expand in size, and personnel mobility is becoming increasingly frequent, posing unprecedented challenges to security management. In these complex environments, ensuring accurate identification of personnel and effective monitoring of their activities is crucial to maintaining public safety and ensuring production order. Efficient and reliable identity recognition technology is the key cornerstone for achieving security management goals.
[0003] Large industrial parks and public spaces often have complex environmental conditions. These areas may be subject to various chemical emissions, widespread dust, and dim lighting, factors that not only interfere with the proper functioning of traditional identification equipment but can also damage it. Large public spaces such as airports and train stations are densely populated, feature complex and changing backgrounds, and are often obstructed by numerous obstructions, presenting significant challenges for traditional identification technology. Furthermore, these locations typically cover a wide area, requiring the coordinated operation of multiple surveillance cameras. However, traditional identification technology struggles with cross-camera tracking and information fusion, making it difficult to achieve continuous and accurate identification and tracking of individuals.
[0004] To solve the above problems, the present invention proposes a long-distance identity recognition method for a large area. Summary of the Invention
[0005] The technical problem to be solved by the present invention is: how to solve the problem that traditional identity recognition technology has shortcomings in cross-camera tracking and information fusion, making it difficult to achieve continuous and accurate identification and trajectory tracking of people. A long-distance identity recognition method for large areas is provided. This method utilizes the unique walking posture of individuals (gait recognition) and multi-target cross-camera tracking technology to achieve identity recognition that does not require active cooperation from people and is suitable for long distances and complex environments, so as to meet the urgent needs of large industrial parks and large public places for efficient, accurate, and intelligent security and personnel management.
[0006] The present invention solves the above technical problems through the following technical solutions, which include the following steps:
[0007] S1: System initialization
[0008] Deploy and configure cameras;
[0009] S2: Video Data Acquisition
[0010] Collect video data in real time, monitor and preprocess the video data quality, and obtain preprocessed image sequences;
[0011] S3: Pedestrian Detection and Segmentation
[0012] The deep learning-based object detection algorithm YOLOv5 is used to detect pedestrians in the preprocessed image sequence, identify pedestrian targets in the image, and output the location information and confidence score of the pedestrian targets. Based on the detected pedestrian target location information, the pedestrian targets are accurately segmented to obtain segmented pedestrian images. The segmented pedestrian images are smoothed to obtain a pedestrian image sequence.
[0013] S4: Gait Recognition
[0014] Gait features are extracted from pedestrian image sequences. The time series data formed by the skeleton joint feature data in the gait features are reconstructed using the gait multi-task recognition model MLDT to obtain the reconstructed skeleton joint time series data. The dynamic features in the gait features are optimized based on the reconstructed skeleton joint time series data to obtain the optimized dynamic features. The optimized dynamic features, the body proportion features in the gait features, and the encoded features processed by the encoder in the gait multi-task recognition model MLDT are fused to obtain a fused identity feature vector. The similarity between the fused identity feature vector and the sample features in the pre-registered feature library is calculated, and the pedestrian identity is determined based on the set similarity threshold. If the similarity exceeds the threshold, the recognition is considered successful and the corresponding identity information is output. If it is lower than the threshold, the identity is marked as unknown and awaits further processing or manual intervention.
[0015] S5: Multi-object Cross-camera Tracking
[0016] Select the target to be tracked, assign it a unique target identifier, record target-related information, track the target within a single camera, perform cross-camera target association processing, and then fuse and optimize the target trajectories obtained by different cameras;
[0017] S6: Result output and application services
[0018] The pedestrian identity information obtained by gait recognition is output to the application service layer in real time, and illegal intrusion warning and personnel trajectory analysis are performed.
[0019] Furthermore, in step S1, when deploying cameras, the cameras are placed at set locations in large industrial parks or large public places so that the camera coverage is comprehensive and has no blind spots; when configuring the cameras, the internal and external parameters of each camera are first determined, and then the acquisition frame rate and image resolution of each camera are set.
[0020] Furthermore, in step S2, when collecting video data, the camera continuously collects video data in the monitored area according to the set collection frame rate and image resolution, and the video data is transmitted in the form of an image sequence; the quality of each frame of the collected image is monitored, and the image clarity index, brightness index, and contrast index are checked. When it is found that the image quality does not meet the index requirements, the image preprocessing operation is triggered; the image preprocessing operation includes image denoising, contrast enhancement, and histogram equalization operations.
[0021] Furthermore, in step S3, an image segmentation algorithm is used to separate the pedestrian target from the background to obtain a binary mask image of the pedestrian target or a pixel-level segmentation result, that is, a pedestrian image is obtained. The image segmentation algorithm includes the SAM algorithm, the U-Net algorithm and the SegNet algorithm.
[0022] Furthermore, in step S4, the extracted gait features include static features and dynamic features, wherein the dynamic features include stride features, stride frequency features, joint angle change sequence features, and body center of gravity movement trajectory features, and the static features include body proportion features and skeleton joint point features;
[0023] The static feature extraction process is as follows:
[0024] S4111: Use a skeleton extraction algorithm to detect the coordinates of pedestrian joints, calculate limb length ratios, height estimates, and body shape indexes, and form static feature vectors that represent body shape, thus obtaining body proportion features.
[0025] S4112: Using a skeleton extraction algorithm to detect the coordinates of the pedestrian's joints, i.e., the coordinates of the skeleton joints, i.e., obtaining the skeleton joint features;
[0026] The process of extracting dynamic features is as follows:
[0027] S4121: Track the motion trajectory of the foot joints using the optical flow method. Divide the individual gait cycles by the periodic changes in the foot motion trajectory. Calculate the displacement between two consecutive foot contact points on the same side as the stride length feature. Count the number of complete gait cycles per unit time as the cadence feature.
[0028] S4122: Use a skeleton extraction algorithm to detect and obtain the coordinates of the skeleton joint points, calculate the joint angles of each joint, and record them to form a time series, that is, obtain a skeleton joint point time series sequence. Based on this time series, extract the mean and variance maximum as time domain statistical features, and simultaneously obtain the main frequency component as frequency domain features through fast Fourier transform. The time domain statistical features and frequency domain features are the joint angle change sequence features;
[0029] S4123: The human body is divided into multiple segments, and a mass ratio is assigned to each segment according to the biomechanical model. The center of mass coordinates of each segment are calculated through the skeleton joints, and the overall center of mass is obtained by weighted averaging the center of mass coordinates of each segment. The overall center of mass is calculated frame by frame, and the center of mass trajectory is smoothed using Kalman filtering. Finally, the swing amplitude and periodicity of the center of mass trajectory are extracted as the characteristics of the body's center of gravity movement trajectory.
[0030] Furthermore, in step S4, the gait multi-task recognition model MLDT includes an encoder, a decoder and a classification head, wherein the encoder maps the noisy or missing skeleton joint time series sequence to a high-dimensional feature space through the linear projection of the embedding layer, and adds the time series position encoding to obtain the encoding features. The encoding features output by the encoder are input into the decoder as global context features. The decoder uses a cross-attention mechanism to align the encoding features with the original input, and gradually reconstructs the complete skeleton joint time series sequence; finally, the classification head performs time series feature aggregation on the reconstructed skeleton joint time series sequence, and outputs the identity probability distribution through the fully connected layer, so as to obtain the fused identity feature vector.
[0031] Furthermore, in the encoder, the specific processing process is as follows:
[0032] S4231: Input standardization and missing information handling
[0033] Input skeleton joint point timing sequence Perform channel normalization, then use bidirectional linear interpolation to fill missing values, and add binary mask marks to joint points below the set confidence level, where T is the time step, N is the number of joint points, and D is the feature dimension;
[0034] S4232: Generation of High-Dimensional Spatiotemporal Embeddings
[0035] Use a learnable linear transformation to map the normalized coordinates to a higher-dimensional space:
[0036]
[0037] in, is the projection parameter, and the output feature dimension is T×N×D;
[0038] S4233: Joint Spatiotemporal Position Coding
[0039] Assign a learnable vector to each time step t via temporal position encoding The periodic timing information of the encoding sequence is encoded, and a learnable vector is assigned to each joint point n through spatial position encoding Encoding the anatomical connection relationship between joints, the final embedding feature is:
[0040]
[0041] S4234: Dynamic Noise Simulation and Feature Enhancement
[0042] During the training phase, random joint masks are applied to the embedded features: some joint features are randomly masked with probability p to simulate the actual noise scene; Gaussian noise ∈~N(0,σ 2 );
[0043] S4235: Encoder input formatting
[0044] Embed features Reshape into a (T×N)×D sequence form, input to the encoder for processing, map to a high-dimensional feature space, and add temporal position encoding to obtain encoding features;
[0045] In the decoder, the specific processing process is as follows:
[0046] S4221: The decoder receives the encoded features output by the encoder and uses a masked multi-head cross attention module to dynamically align the encoded features with the local spatiotemporal information of the original input;
[0047] S4222: Each decoder layer contains a masked cross-attention layer, a multi-head self-attention layer, and a feedforward network. Through residual connections and layer normalization, the global features output by the encoder are fused with the local features obtained by the decoder at each level.
[0048] S4223: The last layer of the encoder maps global features and local features to the original coordinate space through a linear projection layer, and outputs a time series sequence of skeleton joint points
[0049] Furthermore, the specific process of temporal feature aggregation is as follows:
[0050] S431: re-extracting stride features, stride frequency features, joint angle change sequence features, and body center of gravity movement trajectory features from the dynamic features based on the reconstructed skeleton joint point time series to obtain optimized dynamic features;
[0051] S432: Fusing the optimized dynamic features, the body proportion features in the gait features, and the encoded features processed by the encoder in the gait multi-task recognition model MLDT to obtain a fused identity feature vector.
[0052] Furthermore, in step S5, the target related information includes an initial position, a timestamp, a fused identity feature vector, and an appearance feature, which serves as initial state information for target tracking.
[0053] Furthermore, in step S5, the single-camera tracking is performed within the field of view of a single camera using the Mean-Shift Tracking target tracking algorithm to track the target, predict the target position in the next frame image based on the target's motion state, and then perform target search and matching near the predicted position, update the target's position and state information, and thereby simultaneously obtain the target trajectory obtained by single-camera tracking; the cross-camera target association processing is performed when the target enters the field of view of another camera from the field of view of one camera, and uses the ReID calibration algorithm to perform cross-camera association through target features. In the initial stage when the target enters the field of view of the new camera, candidate targets matching the target features are searched in the new camera image, where the target features are the fused identity feature vector and appearance features of the target.
[0054] Furthermore, in step S6, illegal intrusion area rules are set at the application service layer, and illegal intrusion areas are demarcated according to the safety management requirements of industrial parks or public places. When a person is identified as entering the illegal intrusion area, an early warning mechanism is immediately triggered; when performing personnel trajectory analysis, the tracked personnel trajectory is analyzed in real time, and the personnel's residence time, movement speed, and activity range in different areas are calculated to generate a personnel trajectory analysis report to provide data support for park management or public place operations.
[0055] Compared with the existing technology, the present invention has the following advantages: this long-distance identity recognition method for large areas provides an effective solution for large-scene video intelligent security in complex park environments, ensures the accuracy and traceability of personnel identity recognition in high-risk areas, and improves the safety production level of the park; it provides a long-distance multi-target identity recognition solution, realizes real-time monitoring and early warning of illegal personnel intrusion in key protection areas, and provides reliable security protection for public places. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 1 is a flow chart of a method for long-distance identification of a large area according to an embodiment of the present invention;
[0057] Figure 2 1 is a schematic diagram of the implementation process of gait recognition in an embodiment of the present invention. DETAILED DESCRIPTION
[0058] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process. However, the protection scope of the present invention is not limited to the following embodiment.
[0059] like Figure 1 As shown, this embodiment provides a technical solution: a long-distance identity recognition method for a large area, comprising the following steps:
[0060] S1. System initialization
[0061] In this embodiment, step S1 mainly includes the following contents:
[0062] S11. Camera Deployment. Consider the layout of large industrial parks or public spaces and select appropriate locations for HD camera installation. Ensure comprehensive camera coverage with no blind spots, especially in key areas (such as hazardous chemical storage areas in industrial parks and entrances and exits to public spaces) and areas of high human activity (such as aisles in industrial park production halls and waiting halls in public spaces).
[0063] S12. Accurately calibrate each camera to determine its internal parameters (such as focal length, aperture, pixel size, etc.) and external parameters (such as position coordinates, orientation angle, etc.). These parameters will be used for subsequent image correction, target positioning, and trajectory calculation.
[0064] S13. Set the camera's acquisition frame rate, image resolution, and other parameters to balance the amount of data collected with the system's processing capabilities. Also, configure the camera's automatic exposure, white balance, and other functions based on the actual ambient lighting conditions to obtain clear and stable video images.
[0065] S14. System parameter configuration. Initialize the parameters of the gait recognition module, including the parameters of the gait recognition algorithm (such as time step, number of encoder layers, etc.) and the classifier training parameters (such as learning rate, number of iterations, etc.). Adjust these parameters to optimize gait recognition performance based on different application scenarios and accuracy requirements.
[0066] Set parameters for the multi-target cross-camera tracking module, such as the target association similarity threshold, the search range and time interval after target loss, and the parameters of the trajectory detection model. Through experiments and on-site debugging, determine the optimal parameter values to ensure the accuracy and continuity of target tracking.
[0067] Configure the parameters of the data storage and management module, such as storage path, data format, storage period, etc. Establish an efficient data indexing and query mechanism to quickly retrieve and process historical data.
[0068] Start related services of the application service layer, such as identity recognition result query service, illegal intrusion warning service, personnel trajectory analysis service, etc., and set corresponding service ports and permissions.
[0069] S2. Video data acquisition
[0070] In this embodiment, step S2 mainly includes the following contents:
[0071] S21. Real-time data acquisition. The camera continuously collects video data from the monitored area at the set acquisition frame rate and image resolution. The video data is transmitted to the data processing layer as a digital image sequence, ensuring stable and real-time data transmission and preventing data loss or delay.
[0072] S22. Data quality monitoring and preprocessing. Each captured image is quality-monitored, checking image clarity, brightness, contrast, and other indicators. If the image quality does not meet the required indicators (e.g., blurry, too dark, or too bright), image preprocessing is automatically triggered.
[0073] Image preprocessing includes operations such as image denoising, contrast enhancement, and histogram equalization. Appropriate filtering algorithms (such as Gaussian filtering and median filtering) are used to remove noise and improve image clarity. By adjusting the image's brightness and contrast, the difference between the target and the background is enhanced, facilitating subsequent pedestrian detection and segmentation.
[0074] S3. Pedestrian detection and segmentation
[0075] In this embodiment, step S3 mainly includes the following contents:
[0076] S31. Object Detection. Use a deep learning-based object detection algorithm (such as YOLOv5) to detect pedestrians in preprocessed video images. The algorithm uses a pre-trained model to identify pedestrians in the image and output the location information (bounding box coordinates) and confidence score of the pedestrians.
[0077] S32. Pedestrian segmentation. Based on the detected pedestrian location information, accurately segment the pedestrian target. Use an image segmentation algorithm (such as the SAM algorithm, U-Net, SegNet, etc.) to separate the pedestrian from the background and obtain a binary mask image of the pedestrian or a pixel-level segmentation result.
[0078] S33. Process the segmented pedestrian image, smooth the edges, and obtain a more accurate and complete pedestrian outline, providing a clear target object for subsequent gait recognition.
[0079] S4. Gait Recognition
[0080] like Figure 2 As shown, step S4 mainly includes the following contents:
[0081] S41, feature extraction. Gait feature extraction is performed on the segmented pedestrian image sequence (i.e. Figure 2The extracted gait features include dynamic and static features. Motion analysis algorithms (such as optical flow and skeleton extraction algorithms) are used to calculate the pedestrian's joint motion information. Dynamic features such as stride length and frequency, joint angle change sequence, and body center of gravity movement trajectory are obtained through time series analysis. Image processing technology is used to extract static features such as body proportions and skeleton joint points.
[0082] The static feature extraction process is as follows:
[0083] S4111. Skeleton joint feature extraction: Use a skeleton extraction algorithm (such as AlphaPose) to detect the coordinates of pedestrian joints and then obtain skeleton joint features.
[0084] S4112, Body Proportion Feature Extraction: Based on the coordinates of pedestrian joints detected by a skeleton extraction algorithm (such as AlphaPose), the limb length ratio (such as the ratio between shoulder width, torso length, thigh length, etc.), height estimation value and body shape index are calculated to form a static feature vector representing the body shape.
[0085] The process of extracting dynamic features is as follows:
[0086] S4121. Stride and cadence feature extraction: Track the motion trajectory of the foot joints using the optical flow method. Determine individual gait cycles based on periodic changes in the foot motion trajectory (e.g., when the toes are hoeing the ground). Calculate the displacement between two consecutive foot contact points on the same side as the stride feature. Count the number of complete gait cycles per unit time as the cadence feature.
[0087] S4122. Joint Angle Change Sequence Feature Extraction: Use a skeleton extraction algorithm (such as AlphaPose) to detect the coordinates of the skeleton joint points in the pedestrian image, calculate the joint angles of the knee, hip, ankle, and other joints, and record them to form a time series. Based on this sequence, extract time-domain statistical features such as the mean and variance maximum. Simultaneously, obtain frequency-domain features such as the main frequency component through fast Fourier transform. The time-domain statistical features and frequency-domain features are the joint angle change sequence features.
[0088] S4123. Extraction of body center of mass trajectory features. The human body is divided into segments such as the head, trunk, upper limbs, and lower limbs, and each segment is assigned a mass ratio based on the biomechanical model. The center of mass coordinates of each segment are calculated through the skeleton joints (e.g., the center of mass of the trunk is the midpoint between the neck and hip), and the center of mass coordinates of each stage are weighted averaged to obtain the overall center of mass of the target; the overall center of mass is calculated frame by frame, and the center of mass trajectory is smoothed using Kalman filtering; finally, the center of mass trajectory features such as swing amplitude and periodicity are extracted.
[0089] S42. Gait recognition. The existing skeleton-based gait recognition methods inevitably have estimation errors, cross-dislocations, data distortion and other problems in the process of estimating the human skeleton joints. At the same time, the joints are often occluded and self-occluded, resulting in data missing. It is difficult for the existing skeleton-based methods to achieve satisfactory performance in gait recognition. Often, the noise data with incomplete skeleton information is treated as complete information data during the recognition process, and no effective reconstruction and recovery measures are taken to restore the complete gait data. These distorted data and missing data containing noise reduce the performance of the model. For gait identity recognition under noisy data, the present invention adopts a gait multi-task recognition model MLDT based on denoising Transformer to automatically reconstruct the skeleton joint time series data, correct the errors caused by the posture estimation algorithm, and use the temporal smoothing characteristics of the gait trajectory and the self-attention mechanism to encode and embed the repaired skeleton time series data to improve the classification characteristics of the gait embedding feature.
[0090] In this embodiment, the denoising Transformer-based gait multi-task recognition model MLDT adopts an encoder-decoder architecture and a multi-task learning framework, aiming to improve the robustness of the model by jointly optimizing data reconstruction and identity classification tasks.
[0091] The model first receives input of a noisy or missing temporal sequence of skeleton joints, with dimensions T×N×C (corresponding to time steps, number of joints, and coordinate dimensions, respectively). The input data is linearly projected into a high-dimensional feature space via an embedding layer, and temporal position encoding is added to preserve the temporal order of the motion sequence. The encoder consists of multiple layers of stacked Transformer modules, each of which contains a multi-head self-attention mechanism and a feed-forward network (FFN). The multi-head self-attention mechanism captures the spatiotemporal correlations between joints through global dependency modeling, while a masking mechanism is introduced to suppress noise propagation from missing or low-confidence nodes. The encoder output is fed into the decoder as global contextual features. The decoder uses a cross-attention mechanism to align the encoded features with the original input, gradually reconstructing the complete temporal sequence of skeleton joints. The model's multi-task classification head is divided into two parts: the reconstruction branch optimizes data recovery capabilities by calculating the L1 / L2 loss between the reconstructed sequence and the true complete sequence; the identity classification branch performs temporal feature aggregation (such as global average pooling) on the reconstructed skeleton joint point time series sequence, and outputs the identity probability distribution through a fully connected layer. Finally, it jointly optimizes the reconstruction loss and classification loss to achieve high-precision gait recognition under noisy data.
[0092] On this basis, in order to improve the recognition ability and environmental adaptability, in addition to the skeleton joint feature data, the system further enhances the modeling and fusion of the aforementioned gait features (such as stride, cadence, joint angle change sequence, body center of gravity movement trajectory, etc.) based on the repaired (reconstructed) skeleton data. For details, see the following step S43.
[0093] In this embodiment, skeleton joint point data reconstruction is achieved through multi-stage refinement processing of the decoder. The specific processing process is as follows:
[0094] S4221. Feature alignment and masked cross-attention mechanism: The decoder receives the global spatiotemporal features output by the encoder (dimension is T×D, where T is the time step and D is the feature dimension) and the noisy input sequence (T×N×C, where N is the number of joints and C is the coordinate dimension). Random Mask Dropout is introduced after the embedding layer to mask the embedded features of some joints with a certain probability, simulating the noisy scene in the training phase and forcing the model to learn robust representations. A masked multi-head cross-attention module is used to dynamically align the encoded features with the local spatiotemporal information of the original input, where the mask matrix M∈{0,1} T×N Used to mask missing or low-confidence joint points and suppress noise interference.
[0095] S4222, spatiotemporal feature fusion and residual refinement. Each decoder module contains a masked cross-attention layer, a multi-head self-attention layer, and a feedforward network (FFN). Through residual connection (Residual Connection) and layer normalization (LayerNorm), the global features output by the encoder and the local features of the decoder are fused step by step. The formula is expressed as: H l+1 =LayerNorm((H l +FFN(Attention(H l )), where H l is the l-th layer feature.
[0096] S4223, multi-scale reconstruction supervision and output generation. The last layer of the encoder maps the high-dimensional features to the original coordinate space through the linear projection layer and outputs the reconstructed sequence (skeleton joint point time series sequence) The reconstruction loss function is optimized jointly by L1 smoothing loss and speed constraint loss:
[0097]
[0098] Among them, ΔX represents the coordinate difference between adjacent time steps, and λ1 and λ2 are weight coefficients.
[0099] In this embodiment, the encoding embedding of skeleton time series data is achieved through spatiotemporal feature mapping and noise robustness enhancement. The specific steps are as follows:
[0100] S4231, input normalization and missing processing. Input skeleton sequence Perform channel normalization, the calculation formula is: Where μ and σ are the mean and standard deviation of the joint coordinates in the training set. Missing values are filled using bidirectional linear interpolation, and low-confidence joint points are marked with binary masks.
[0101] S4232, High-dimensional spatiotemporal embedding generation. Use a learnable linear transformation to map normalized coordinates to high-dimensional space:
[0102]
[0103] in, is the projection parameter, and the output feature dimension is T×N×D.
[0104] S4233, Joint spatiotemporal position coding. Temporal position coding: assign a learnable vector to each time step t Encodes the periodic timing information of the motion sequence. Spatial position encoding: assigns a learnable vector to each joint point n Encode the anatomical connection relationship between joints. The final embedding feature is:
[0105]
[0106] S4234, dynamic noise simulation and feature enhancement. During the training phase, random joint masking is applied to the embedded features: some joint features are randomly masked with probability p to simulate the actual noise scene. Feature jittering technology is used to add Gaussian noise ∈~N(0,σ 2 ), improving the robustness of the model.
[0107] S4235, encoder input formatting. Embedding features The data is reshaped into a (T×N)×D sequence and fed into the encoder for spatiotemporal dependency modeling. The encoder output undergoes feature aggregation (e.g., global average pooling) to obtain a T×D-dimensional global context feature for subsequent identity determination and feature comparison.
[0108] S43. Fusion feature extraction and gait representation optimization.
[0109] To further improve the accuracy and robustness of identity recognition, after completing the denoising and reconstruction of the skeleton joint time series data, the system re-extracts, modifies, and optimizes the modeling of some of the gait features obtained in the preliminary feature extraction phase based on the reconstructed skeleton joint time series data. In this phase, the system no longer repeats the preliminary extraction process, but instead achieves high-quality feature fusion and final gait representation generation through the following steps:
[0110] S431. Reconstruction data-driven stride length and cadence extraction: Based on the restored foot joint motion trajectory, stride length and cadence are recalculated based on more accurate gait cycle recognition to enhance the credibility of rhythm information;
[0111] S432, Joint Angle Sequence Feature Optimization Extraction: Using the complete skeleton sequence, recalculate the knee, ankle, hip and other joint angle change curves, and extract their statistical and frequency domain features to capture individual motion control characteristics;
[0112] S433, Body Center of Gravity Trajectory Correction and Modeling: By accurately estimating the center of mass of skeleton points, a smoother and more stable body center of gravity trajectory sequence is obtained, and its swing pattern and periodic structure are calculated to enhance the ability to depict gait style.
[0113] S434, Feature Fusion and Encoding: The restored and optimized multi-dimensional gait features, the initially extracted body proportion features, and the encoded features generated by the skeleton embedding module in the MLDT model (used to implement steps S4231 to S4235) are jointly modeled. Fusion methods include concatenation, weighted averaging, and modal attention mechanisms to form a unified, highly discriminative gait feature representation.
[0114] S44, the identity determination module realizes open set recognition of pedestrian identities through feature comparison and dynamic threshold decision making. The specific process is as follows:
[0115] S441, based on the fusion identity feature vector obtained in S434 Pre-registered signature database where y i For known identity tags, many-to-one similarity calculation is performed and the cosine similarity is used to measure the correlation between features. The formula is:
[0116]
[0117] S442, determine based on the preset similarity threshold θ: if there is s i ≥θ, then it is determined to be the corresponding identity y i And output the result; if all s i<θ, it is marked as "unknown identity", triggering an early warning signal and starting the manual review process. To adapt to different scenario requirements, the threshold θ supports dynamic adjustment;
[0118] S5. Multi-target cross-camera tracking
[0119] In this embodiment, step S5 mainly includes the following contents:
[0120] S51. Target Initialization and Identity Assignment. A target is selected and assigned a unique target identity. Information such as its initial position, timestamp, fused identity feature vector, and appearance features (such as clothing color and texture) are recorded as the initial state for target tracking.
[0121] S52, Single-Camera Tracking. A pedestrian is tracked within the field of view of a single camera using the Mean-Shift Tracking algorithm. The target's position in the next frame is predicted based on the pedestrian's motion state. Target search and matching is then performed near the predicted position, updating the target's position and state information.
[0122] S53. Cross-camera object association. When an object enters the field of view of one camera from another, a ReID calibration algorithm is used to perform cross-camera association based on object features (such as fused identity feature vectors and appearance features). During the initial phase of the object entering the field of view of a new camera, candidate objects matching the object features are searched in the new camera image.
[0123] S54. Trajectory fusion and optimization. Fusion and optimization are performed on the target trajectories obtained by different cameras. A trajectory splicing algorithm is used to connect the segmented trajectories into a complete target trajectory. The target trajectory is also smoothed to remove noise and jitter, improving trajectory accuracy and continuity.
[0124] S6. Result output and application services
[0125] S61. Output of identification results. The pedestrian's identity information obtained from gait recognition is output to the application service layer in real time. The identity information includes details such as name (if registered in the system), identification number, identification time, and identification location (camera number and location).
[0126] For pedestrians whose identities cannot be identified, their fused identity feature vector, appearance features, and related image information are output for subsequent manual investigation or further analysis and processing.
[0127] S62, illegal intrusion warning. Set illegal intrusion area rules at the application service layer, and demarcate specific restricted areas (such as dangerous areas in industrial parks, backend management areas in public places, etc.) according to the safety management requirements of industrial parks or public places. When the system identifies a pedestrian entering an illegal intrusion area, it immediately triggers the warning mechanism. The warning information is sent to security personnel through various methods such as sound alarms, SMS notifications, system pop-ups, etc., and the location and trajectory of the intrusion target are highlighted on the monitoring interface to facilitate a quick response by security personnel.
[0128] S63. Personnel trajectory analysis. Real-time analysis of tracked personnel trajectories is performed to calculate statistical information such as the duration of stay, movement speed, and range of activity in different areas. Personnel trajectory analysis reports are generated to provide data support for park management or public space operations, helping to optimize personnel layout, resource allocation, and safety management strategies.
[0129] Supports historical track query function, security personnel or management personnel can query the historical track of specific personnel based on time, location, personnel identity and other conditions, in order to trace the event process, investigate abnormal behavior, etc.
[0130] In summary, the long-distance identity recognition method for large areas in the above-mentioned embodiment provides an effective solution for large-scale video intelligent security in complex campus environments, ensures the accuracy and traceability of personnel identity recognition in high-risk areas, and improves the level of safe production in the park; provides a long-distance multi-target identity recognition solution, realizes real-time monitoring and early warning of illegal personnel intrusion in key protection areas, and provides reliable security protection for public places.
[0131] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A long-distance identity recognition method for a large area, characterized in that: The following steps are involved: S1: System initialization Deploy and configure cameras; S2: Video Data Acquisition Collect video data in real time, monitor and preprocess the video data quality, and obtain preprocessed image sequences; S3: Pedestrian Detection and Segmentation Use the deep learning-based object detection algorithm YOLOv5 to detect pedestrians in the preprocessed image sequence, identify pedestrian targets in the image, and output the location information and confidence score of the pedestrian targets; Based on the detected pedestrian target position information, the pedestrian target is accurately segmented to obtain the segmented pedestrian image; the segmented pedestrian image is smoothed to obtain a pedestrian image sequence; S4: Gait Recognition Gait features are extracted from pedestrian image sequences. The time series data formed by the skeleton joint feature data in the gait features are reconstructed using the gait multi-task recognition model MLDT to obtain the reconstructed skeleton joint time series data. The dynamic features in the gait features are optimized based on the reconstructed skeleton joint time series data to obtain the optimized dynamic features. The optimized dynamic features, the body proportion features in the gait features, and the encoded features processed by the encoder in the gait multi-task recognition model MLDT are fused to obtain a fused identity feature vector. The similarity between the fused identity feature vector and the sample features in the pre-registered feature library is calculated, and the pedestrian identity is determined based on the set similarity threshold. If the similarity exceeds the threshold, the recognition is considered successful and the corresponding identity information is output. If it is lower than the threshold, the identity is marked as unknown and awaits further processing or manual intervention. S5: Multi-object Cross-camera Tracking Select the target to be tracked, assign it a unique target identifier, record target-related information, track the target within a single camera, perform cross-camera target association processing, and then fuse and optimize the target trajectories obtained by different cameras; S6: Result output and application services The pedestrian identity information obtained by gait recognition is output to the application service layer in real time, and illegal intrusion warning and personnel trajectory analysis are performed.
2. A long-distance identity recognition method for a large area according to claim 1, characterized in that: In step S1, when deploying cameras, the cameras are placed at set locations in large industrial parks or large public places so that the camera coverage is comprehensive and has no blind spots; when configuring the cameras, the internal and external parameters of each camera are first determined, and then the acquisition frame rate and image resolution of each camera are set.
3. The method for long-distance identification of a large area according to claim 1, characterized in that: In step S2, when collecting video data, the camera continuously collects video data in the monitoring area according to the set acquisition frame rate and image resolution, and the video data is transmitted in the form of an image sequence; the quality of each frame of the collected image is monitored, and the image clarity index, brightness index, and contrast index are checked. When it is found that the image quality does not meet the index requirements, the image preprocessing operation is triggered; the image preprocessing operation includes image denoising, contrast enhancement, and histogram equalization operations.
4. The method for long-distance identification of a large area according to claim 1, characterized in that: In step S3, an image segmentation algorithm is used to separate the pedestrian target from the background to obtain a binary mask image of the pedestrian target or a pixel-level segmentation result, that is, a pedestrian image. The image segmentation algorithm includes the SAM algorithm, the U-Net algorithm and the SegNet algorithm.
5. The method for long-distance identification of a large area according to claim 4, characterized in that: In step S4, the extracted gait features include static features and dynamic features, wherein the dynamic features include stride features, cadence features, joint angle change sequence features, and body center of gravity movement trajectory features, and the static features include body proportion features and skeleton joint point features; The static feature extraction process is as follows: S4111: Use a skeleton extraction algorithm to detect the coordinates of pedestrian joints, calculate limb length ratios, height estimates, and body shape indexes, and form static feature vectors that represent body shape, thus obtaining body proportion features. S4112: Using a skeleton extraction algorithm to detect the coordinates of the pedestrian's joints, i.e., the coordinates of the skeleton joints, i.e., obtaining the skeleton joint features; The process of extracting dynamic features is as follows: S4121: Track the motion trajectory of the foot joints using the optical flow method. Divide the individual gait cycles by the periodic changes in the foot motion trajectory. Calculate the displacement between two consecutive foot contact points on the same side as the stride length feature. Count the number of complete gait cycles per unit time as the cadence feature. S4122: Use a skeleton extraction algorithm to detect and obtain the coordinates of the skeleton joint points, calculate the joint angles of each joint, and record them to form a time series, that is, obtain a skeleton joint point time series sequence. Based on this time series, extract the mean and variance maximum as time domain statistical features, and simultaneously obtain the main frequency component as frequency domain features through fast Fourier transform. The time domain statistical features and frequency domain features are the joint angle change sequence features; S4123: The human body is divided into multiple segments, and a mass ratio is assigned to each segment according to the biomechanical model. The center of mass coordinates of each segment are calculated through the skeleton joints, and the overall center of mass is obtained by weighted averaging the center of mass coordinates of each segment. The overall center of mass is calculated frame by frame, and the center of mass trajectory is smoothed using Kalman filtering. Finally, the swing amplitude and periodicity of the center of mass trajectory are extracted as the characteristics of the body's center of gravity movement trajectory.
6. The method for long-distance identification of a large area according to claim 5, characterized in that: In step S4, the gait multi-task recognition model MLDT includes an encoder, a decoder and a classification head, wherein the encoder maps the noisy or missing skeleton joint time series sequence to a high-dimensional feature space through the linear projection of the embedding layer, and adds the time series position encoding to obtain the encoding features. The encoding features output by the encoder are input into the decoder as global context features. The decoder uses a cross-attention mechanism to align the encoding features with the original input, and gradually reconstructs the complete skeleton joint time series sequence; finally, the classification head aggregates the time series features of the reconstructed skeleton joint time series sequence, and outputs the identity probability distribution through the fully connected layer, so as to obtain the fused identity feature vector.
7. The method for long-distance identification of a large area according to claim 6, characterized in that: In the encoder, the specific processing process is as follows: S4231: Input standardization and missing information handling Input skeleton joint point timing sequence Perform channel normalization, then use bidirectional linear interpolation to fill missing values, and add binary mask marks to joint points below the set confidence level, where T is the time step, N is the number of joint points, and D is the feature dimension; S4232: Generation of High-Dimensional Spatiotemporal Embeddings Use a learnable linear transformation to map the normalized coordinates to a higher-dimensional space: in, is the projection parameter, and the output feature dimension is T×N×D; S4233: Joint Spatiotemporal Position Coding Assign a learnable vector to each time step t via temporal position encoding The periodic timing information of the encoding sequence is encoded, and a learnable vector is assigned to each joint point n through spatial position encoding Encoding the anatomical connection relationship between joints, the final embedding feature is: S4234: Dynamic Noise Simulation and Feature Enhancement During the training phase, random joint masks are applied to the embedded features: some joint features are randomly masked with probability p to simulate the actual noise scene; Gaussian noise ∈~N(0,σ 2 ); S4235: Encoder input formatting Embed features Reshape into a (T×N)×D sequence form, input to the encoder for processing, map to a high-dimensional feature space, and add temporal position encoding to obtain encoding features; In the decoder, the specific processing process is as follows: S4221: The decoder receives the encoded features output by the encoder and uses a masked multi-head cross attention module to dynamically align the encoded features with the local spatiotemporal information of the original input; S4222: Each decoder layer contains a masked cross-attention layer, a multi-head self-attention layer, and a feedforward network. Through residual connections and layer normalization, the global features output by the encoder are fused with the local features obtained by the decoder at each level. S4223: The last layer of the encoder maps global features and local features to the original coordinate space through a linear projection layer, and outputs a time series sequence of skeleton joint points 8. The method for long-distance identification of a large area according to claim 6, characterized in that: The specific process of temporal feature aggregation is as follows: S431: re-extracting stride features, stride frequency features, joint angle change sequence features, and body center of gravity movement trajectory features from the dynamic features based on the reconstructed skeleton joint point time series to obtain optimized dynamic features; S432: Fusing the optimized dynamic features, the body proportion features in the gait features, and the encoded features processed by the encoder in the gait multi-task recognition model MLDT to obtain a fused identity feature vector.
9. The method for long-distance identification of a large area according to claim 8, characterized in that: In step S5, the target related information includes an initial position, a timestamp, a fused identity feature vector, and an appearance feature, which serves as initial state information for target tracking.
10. The method for long-distance identification of a large area according to claim 9, characterized in that: In step S5, single-camera tracking is to track the target within the field of view of a single camera using the Mean-Shift Tracking target tracking algorithm, predict the target position in the next frame image based on the target's motion state, then perform target search and matching near the predicted position, update the target's position and state information, and simultaneously obtain the target trajectory obtained by single-camera tracking; Cross-camera target association processing means that when a target enters the field of view of one camera from another, the ReID calibration algorithm is used to perform cross-camera association through target features. In the initial stage when the target enters the field of view of the new camera, candidate targets matching the target features are searched in the new camera image, where the target features are the fused identity feature vector and appearance features of the target.
Citation Information
Patent Citations
Overlapped domain dual-camera target tracking system and method
CN103997624A
Multi-mode pedestrian identity recognition method and system based on pedestrian appearance and gait information
CN111860291A
Cited By
Deep learning-based medical self-service check-in terminal interactor identity binding method and system
CN121389094A
Campus security monitoring method based on image recognition technology
CN121486534A