Intelligent elevator personnel identification method based on deep learning

By fusing static features and dynamic temporal information through a dual-branch deep network model, the problem of recognizing complex posture changes and dynamic interactive behaviors in traditional elevator recognition methods is solved, thereby improving the accuracy and stability of personnel recognition in elevator environments.

CN120954055APending Publication Date: 2025-11-14GUANGXI OURI ELEVATOR SERVICE GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511104019.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional intelligent elevator personnel recognition methods struggle to capture complex posture changes and dynamic interactive behaviors. They are susceptible to changes in lighting, occlusion interference, and perspective distortion. They are unable to model temporal action patterns and spatial joint relationships, resulting in high false detection rates in scenarios with many people, rapid movement, or partial occlusion. Furthermore, they lack the ability to recognize short-lived actions, continuous behaviors, and motion-blurred scenes, leading to a high misjudgment rate.

Method used

A dual-branch deep network model is adopted, which combines static features and dynamic temporal information. It models human behavior and spatial distribution through multimodal data, uses optical flow algorithm to quantify motion trajectory, and combines recurrent network to model long-term temporal dependencies to construct a dual-branch deep network model for personnel identification.

Benefits of technology

It effectively solves the problem of decreased recognition accuracy caused by changes in viewing angle, people occlusion and dynamic blur in elevator environments, enhances stability for fast-moving and blurred frames, and reduces the false judgment rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954055A_ABST
    Figure CN120954055A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent elevator personnel identification method based on deep learning. The method comprises the steps of data collection in an elevator, data preprocessing, personnel identification model construction and elevator personnel identification. The invention relates to the technical field of intelligent personnel identification, in particular to an intelligent elevator personnel identification method based on deep learning, which comprises the following steps: acquiring data in an elevator to obtain personnel identification original data; a data preprocessing method of video data processing, personnel posture data processing, optical flow data optimization and data set segmentation is adopted; a double-branch deep network model is adopted as a personnel identification model, so that the problem that the identification precision is reduced due to view angle change, personnel shielding, dynamic fuzziness and the like in an elevator environment is effectively solved; a method of capturing continuous inter-frame motion characteristics through time sequence modeling by adopting motion branches is adopted, real personnel actions and environment interference are effectively distinguished, and the stability of fast moving and fuzzy frames is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent personnel recognition technology, specifically to an intelligent elevator personnel recognition method based on deep learning. Background Technology

[0002] Intelligent elevator personnel identification uses technologies such as deep learning and computer vision to automatically identify people inside elevators. It collects information from inside the elevator through cameras or sensors, and performs real-time detection of the presence of people inside the elevator and abnormal behavior. This ensures the safe operation of elevators, improves elevator management efficiency, and enables more effective provision of intelligent elevator services.

[0003] However, traditional intelligent elevator personnel recognition methods suffer from several technical problems. They struggle to capture complex posture changes and dynamic interactions, are susceptible to lighting variations, occlusion interference, and viewpoint distortion, and cannot model temporal action patterns and spatial joint relationships. This results in high false detection rates in scenarios with dense crowds, rapid movement, or partial occlusion. Furthermore, traditional intelligent elevator personnel recognition methods rely solely on single-frame image analysis or simple inter-frame differencing, which is insufficient for recognizing brief actions, continuous behaviors, and motion-blurred scenes. When people move rapidly, image blurring can cause feature distortion, making it difficult to construct a coherent behavioral logic chain, ultimately leading to increased false detection rates in dynamic scenes. Summary of the Invention

[0004] To address the above issues and overcome the shortcomings of existing technologies, this invention provides a deep learning-based intelligent elevator personnel recognition method. Traditional intelligent elevator personnel recognition methods struggle to capture complex posture changes and dynamic interactions, are susceptible to lighting variations, occlusion interference, and viewpoint distortion, and cannot model temporal action patterns and spatial joint relationships, leading to high false detection rates in crowded, fast-moving, or partially occluded scenarios. This solution creatively employs a dual-branch deep network model as the personnel recognition model. By fusing static features and dynamic temporal information, and combining multimodal data to model human behavior and spatial distribution in complex scenarios, it effectively solves the problems caused by viewpoint changes in elevator environments. This solution addresses the issue of decreased recognition accuracy caused by personnel occlusion and dynamic blur. Traditional intelligent elevator personnel recognition methods rely solely on single-frame image analysis or simple inter-frame differencing, resulting in insufficient recognition capabilities for short-lived actions, continuous behaviors, and motion-blurred scenes. Furthermore, these methods are prone to feature distortion due to image blurring during rapid personnel movement, making it difficult to construct a coherent behavioral logic chain and ultimately leading to increased misjudgment rates in dynamic scenes. This solution creatively employs a motion branching approach to capture continuous inter-frame motion features through temporal modeling. It utilizes optical flow algorithms to quantize motion trajectories and combines this with recurrent networks to model long-term temporal dependencies, effectively distinguishing between real personnel actions and environmental interference, thus enhancing stability for fast-moving and blurred frames.

[0005] The technical solution adopted by this invention is as follows: This invention provides a deep learning-based intelligent elevator personnel recognition method, which includes the following steps:

[0006] Step S1: Data collection inside the elevator;

[0007] Step S2: Data preprocessing;

[0008] Step S3: Personnel identification model construction;

[0009] Step S4: Elevator personnel identification.

[0010] Further, in step S1, the elevator data acquisition is used to collect the raw data required for intelligent elevator personnel identification. Specifically, by performing data acquisition, a raw personnel identification dataset is obtained. The raw personnel identification dataset specifically includes a historical raw identification dataset and a current raw identification dataset. Both the historical raw identification dataset and the current raw identification dataset include joint coordinate data collected by the depth camera and elevator monitoring video data. The historical raw identification dataset also includes personnel posture annotation data, personnel number annotation data, and personnel occlusion annotation data.

[0011] Further, in step S2, the data preprocessing is used to preprocess the collected raw data, specifically including the following steps:

[0012] Step S21: Video data processing, used to eliminate video image distortion, ensure video temporal consistency and focus video key frames, specifically to perform geometric distortion correction, time synchronization alignment and key frame extraction on the elevator monitoring video data;

[0013] Step S22: Personnel posture data processing, used to process the joint coordinate data collected by the depth camera. Specifically, manual quality inspection is used to correct the labeling error of the joint coordinate data collected by the depth camera, and the bone length is normalized according to the national standard.

[0014] Step S23: Optical flow data optimization, used to extract the precise motion vectors of people in the elevator scene. Specifically, it restores the image blur caused by the rapid movement of people by using a blurred Gaussian kernel function and fast Fourier transform, and filters out noise by median filtering and direction consistency detection.

[0015] Step S24: Dataset segmentation, used to segment the historical dataset into a training set and a test set, specifically, to segment the original historical recognition dataset into a recognition training set and a recognition test set;

[0016] Step S25: Perform preprocessing, specifically by preprocessing the current recognition raw dataset through video data processing, personnel posture data processing, and optical flow data optimization to obtain the dataset to be recognized; and by preprocessing the historical recognition raw dataset through video data processing, personnel posture data processing, optical flow data optimization, and dataset segmentation to obtain the recognition training set and the recognition test set.

[0017] Further, in step S3, the personnel recognition model construction is used to construct the model required for personnel recognition in the intelligent elevator, specifically to construct a dual-branch deep network model as the personnel recognition model. The dual-branch deep network model specifically includes a depth correction module, a visible light branch, a motion branch, a graph convolution module, and a multi-dimensional output module.

[0018] The construction of the personnel identification model specifically includes the following steps:

[0019] Step S31: Construction of the depth correction module, used to further eliminate image distortion caused by installation angle and lens characteristics. Specifically, it involves constructing a multi-scale feature pyramid to analyze the image distortion pattern, generating a pixel-level deformation correction field, and obtaining a geometrically corrected image that conforms to the real spatial distribution. The steps include:

[0020] Step S311: Multi-scale distortion analysis is used to capture the distortion features of different scales generated by the camera in the elevator scene. Specifically, the image pyramid features are extracted step by step through a five-layer coding structure, and the receptive field is expanded by dilated convolution to obtain the initial deformation field.

[0021] Step S312: Deformation field regression, used to map abstract features into interpretable geometric correction parameters, specifically by fusing shallow details and deep semantic features through deconvolutional layers and skip connections to obtain an interpretable offset field;

[0022] Step S313: Adaptive remapping, used to achieve high-precision image geometric correction, specifically by calculating the weighted pixel values ​​of the non-integer coordinate neighborhood in the input data, and combining the Gaussian kernel function to suppress interpolation noise, to obtain the correction result;

[0023] Step S32: Visible light branch construction, used to capture the static features of people in the elevator scene. Specifically, it improves the depthwise separable convolutional structure of MobileNetV3, reduces computational cost while using a differentiable pruning activation function, and combines adaptive pooling to statistically determine channel importance. The steps include:

[0024] Step S321: Design adaptive pooling to focus on the effective human body region in a noisy elevator scene. Specifically, channel importance is statistically analyzed using global average pooling, and channel attention weights are generated using Softmax. The formula used is as follows:

[0025] ;

[0026] In the formula, This represents the adaptive pooling function. This represents a function that performs global average pooling down to a resolution of c×c. represents the learnable weight matrix, and x represents the input independent variable;

[0027] Step S322: Design an improved MobileNetV3 model, specifically by replacing the ReLU activation function with a differentiable pruning activation function in the feature calculation of each layer, and combining it with adaptive pooling;

[0028] Step S323: Output layer design, specifically, using depthwise separable convolution combined with pointwise convolution to obtain visible light branch output features;

[0029] Step S33: Motion branch construction, used to capture the dynamic behavior patterns of people in the elevator scene. Specifically, it calculates the motion vectors of adjacent frames using the total variational-L1 optical flow algorithm, inputs them into the gated cyclic unit to model the temporal dependency, and obtains the aggregated motion features of 5 consecutive frames. The formula used is as follows:

[0030] ;

[0031] In the formula, Represents the motion vector at time t. Represents the total variational -L1 optical flow function. This represents the corrected data at time t. Indicates the characteristics of aggregation motion. This represents a gated loop unit function. This represents the motion vector at time t-2. This represents the motion vector at time t+2;

[0032] Step S34: Construction of the graph convolution module, used to integrate visible light branch output features and aggregated motion features, and to construct a human joint relationship model that conforms to biomechanical principles based on graph convolution. Specifically, visible light branch output features and aggregated motion features are fused through dynamic gating weights, and the positions of major human joints are located through key point heatmaps. After graph convolution propagation, human perception features are obtained. The steps include:

[0033] Step S341: Dynamic gating feature fusion, specifically, compressing the feature channel dimension through global average pooling, generating dynamic gating weights using a two-layer fully connected network, and weighted fusing the visible light branch output features and aggregated motion features;

[0034] Step S342: Generating a key point heatmap. The key points correspond to human body key points in COCO format. The formula used to generate the key point heatmap is as follows:

[0035] ;

[0036] In the formula, Represents a heatmap of key points. This represents a convolution function with a kernel size of 1×1. Indicates fusion characteristics;

[0037] Step S343: Joint relationship modeling, specifically, modeling the nonlinear relationships between joints using a learnable adjacency matrix, with the following formula:

[0038] ;

[0039] In the formula, Represents the elements of the adjacency matrix. This represents the LeakyReLU activation function. This represents the learnable attention weight matrix. This represents the heatmap of the m-th key point. This represents the heatmap of the nth key point. This represents the heatmap of the k-th key point;

[0040] Step S344: Graph convolution propagation;

[0041] Step S35: Construct a multi-dimensional output module for synchronously outputting personnel behavior recognition, personnel quantity, and occlusion status. The steps include:

[0042] Step S351: Output of personnel behavior recognition, specifically by obtaining the personnel behavior recognition result through a fully connected layer based on the softmax function;

[0043] Step S352: Output the number of people. Specifically, the key point heatmap is normalized by Softmax and then linearly mapped to a density map to obtain the number of people.

[0044] Step S353: Occlusion state output, specifically by performing max pooling on human perception features and then classifying them using a multilayer perceptron to obtain the occlusion state result;

[0045] Step S36: Construct and train the model. Specifically, construct a dual-branch deep network model by constructing the depth correction module, the visible light branch, the motion branch, the graph convolution module, and the multidimensional output module. Train the model based on the recognition training set and verify the model performance based on the recognition test set to obtain the dual-branch deep network model as the person recognition model.

[0046] Further, in step S4, the elevator personnel identification specifically involves using the dataset to be identified as the input to the personnel identification model to perform intelligent elevator personnel identification, obtaining elevator personnel identification reference results, and comprehensively evaluating the personnel situation inside the elevator based on the elevator personnel identification reference results. The elevator personnel identification reference results specifically include personnel behavior identification results, personnel quantity results, and occlusion status results.

[0047] The beneficial effects achieved by the present invention using the above solution are as follows:

[0048] (1) In view of the technical problems of traditional intelligent elevator personnel recognition methods, such as difficulty in capturing complex posture changes and dynamic interaction behaviors, susceptibility to changes in lighting, occlusion interference and view distortion, and inability to model temporal action patterns and spatial joint associations, resulting in high false detection rates in scenarios with dense crowds, rapid movement or partial occlusion, this solution creatively adopts a dual-branch deep network model as the personnel recognition model. By integrating static features and dynamic temporal information, and combining multimodal data to model human behavior and spatial distribution in complex scenarios, it effectively solves the problem of decreased recognition accuracy caused by view changes, personnel occlusion and dynamic blur in the elevator environment.

[0049] (2) In view of the technical problems of traditional intelligent elevator personnel recognition methods, which rely only on single-frame image analysis or simple inter-frame difference, have insufficient recognition ability for short-term actions, continuous behaviors and motion-blurred scenes, and are prone to feature distortion due to image blur when people move quickly, making it difficult to construct a coherent behavioral logic chain, which ultimately leads to an increased misjudgment rate in dynamic scenes, this solution creatively adopts the method of capturing continuous inter-frame motion features through motion branching and temporal modeling, uses optical flow algorithm to quantize motion trajectory, and combines recurrent network to model long temporal dependencies, effectively distinguishing real personnel actions from environmental interference, and enhancing the stability of fast-moving and blurry frames. Attached Figure Description

[0050] Figure 1 A flowchart illustrating a deep learning-based intelligent elevator personnel recognition method provided by this invention;

[0051] Figure 2 This is a flowchart illustrating the data preprocessing process in step S2.

[0052] Figure 3 A flowchart illustrating the process of building the personnel identification model in step S3;

[0053] Figure 4 This is a flowchart illustrating the construction of the convolution module in step S34.

[0054] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation

[0055] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0056] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0057] Example 1, see Figure 1 The technical solution adopted by this invention is as follows: This invention provides a deep learning-based intelligent elevator personnel recognition method, which includes the following steps:

[0058] Step S1: Data collection inside the elevator;

[0059] Step S2: Data preprocessing;

[0060] Step S3: Personnel identification model construction;

[0061] Step S4: Elevator personnel identification.

[0062] Example 2, see Figure 1 In step S1, the elevator data acquisition is used to collect the raw data required for intelligent elevator personnel identification. Specifically, it involves acquiring a raw personnel identification dataset by performing data acquisition. The raw personnel identification dataset includes a historical raw identification dataset and a current raw identification dataset. Both the historical raw identification dataset and the current raw identification dataset include key point coordinate data and elevator monitoring video data collected by a depth camera. The historical raw identification dataset also includes personnel posture annotation data, personnel quantity annotation data, and personnel occlusion annotation data. The elevator monitoring video data in the historical raw identification dataset specifically includes elevator monitoring video data from different time periods, different passenger numbers, different occlusion conditions, different elevator types, and different elevator camera angles.

[0063] Example 3, see Figure 1 and Figure 2 This embodiment is based on the above embodiment. In step S2, the data preprocessing is used to preprocess the collected raw data, specifically including the following steps:

[0064] Step S21: Video data processing, used to eliminate video image distortion, ensure video temporal consistency and focus video key frames, specifically to perform geometric distortion correction, time synchronization alignment and key frame extraction on the elevator monitoring video data;

[0065] The geometric distortion correction is used to address the distortion problem caused by lens distortion. Specifically, it corrects the video data frame by frame by reading the camera intrinsic parameter matrix and distortion coefficients and applying perspective transformation and interpolation algorithms.

[0066] The time synchronization alignment is used to ensure the timing synchronization of video data, specifically by calibrating the device clock through the Network Time Protocol and compensating for the transmission delay of video acquisition.

[0067] The keyframe extraction is used to reduce data redundancy and focus on effective segments. Specifically, it is to filter out keyframes of personnel entering and exiting and personnel posture changes based on changes in elevator door status, elevator load, and motion ambiguity.

[0068] Step S22: Personnel posture data processing, used to process the joint coordinate data collected by the depth camera. Specifically, manual quality inspection is used to correct the labeling error of the joint coordinate data collected by the depth camera, and the bone length is normalized according to the national standard.

[0069] Step S23: Optical flow data optimization, used to extract the precise motion vectors of people in the elevator scene. Specifically, it restores the image blur caused by the rapid movement of people by using a blurred Gaussian kernel function and fast Fourier transform, and filters out noise by median filtering and direction consistency detection.

[0070] Step S24: Dataset segmentation, used to segment the historical dataset into a training set and a test set, specifically, to segment the original historical recognition dataset into a recognition training set and a recognition test set;

[0071] Step S25: Perform preprocessing, specifically by preprocessing the current recognition raw dataset through video data processing, personnel posture data processing, and optical flow data optimization to obtain the dataset to be recognized; and by preprocessing the historical recognition raw dataset through video data processing, personnel posture data processing, optical flow data optimization, and dataset segmentation to obtain the recognition training set and the recognition test set.

[0072] Example 4, see Figure 1 , Figure 3 and Figure 4 This embodiment is based on the above embodiment. In step S3, the personnel recognition model is constructed to build the model required for personnel recognition in the intelligent elevator. Specifically, a dual-branch deep network model is constructed as the personnel recognition model. The dual-branch deep network model specifically includes a depth correction module, a visible light branch, a motion branch, a graph convolution module, and a multi-dimensional output module.

[0073] The construction of the personnel identification model specifically includes the following steps:

[0074] Step S31: Construction of the depth correction module, used to further eliminate image distortion caused by installation angle and lens characteristics. Specifically, it involves constructing a multi-scale feature pyramid to analyze the image distortion pattern, generating a pixel-level deformation correction field, and obtaining a geometrically corrected image that conforms to the real spatial distribution. The steps include:

[0075] Step S311: Multi-scale distortion analysis, used to capture the distortion features at different scales produced by the camera in the elevator scene. Specifically, it extracts the image pyramid features step by step through a five-layer coding structure, and combines it with dilated convolution to expand the receptive field to obtain the initial deformation field. The formula used is as follows:

[0076] ;

[0077] In the formula, This represents the deformation field characteristics of the l-th layer. Represents the convolution function. Indicates input data, Represents the max pooling function. This represents the deformation field characteristics of the (l-1)th layer. This represents dilated convolution, with the dilation rate set to 1, 2, and 4 to gradually increase the receptive field;

[0078] Step S312: Deformation field regression, used to map abstract features into interpretable geometric correction parameters, specifically by fusing shallow details and deep semantic features through deconvolutional layers and skip connections to obtain an interpretable offset field;

[0079] Step S313: Adaptive remapping, used to achieve high-precision image geometric correction, specifically involves calculating the weighted pixel values ​​of the non-integer coordinate neighborhoods in the input data, combining this with a Gaussian kernel function to suppress interpolation noise, and obtaining the correction result. The formula used is as follows:

[0080] ;

[0081] In the formula, This represents the value of the correction data at coordinates (u, v). This indicates that the input data is in coordinates. The value, This represents the Gaussian kernel function with a standard deviation of 0.8. Indicates coordinates pixel values, This represents the pixel value at coordinates (u, v). Indicates the offset field in coordinates The value of ds represents the downsampling factor;

[0082] Step S32: Visible light branch construction, used to capture the static features of people in the elevator scene. Specifically, it improves the depthwise separable convolutional structure of MobileNetV3, reduces computational cost while using a differentiable pruning activation function, and combines adaptive pooling to statistically determine channel importance. The steps include:

[0083] Step S321: Design adaptive pooling to focus on the effective human body region in a noisy elevator scene. Specifically, channel importance is statistically analyzed using global average pooling, and channel attention weights are generated using Softmax. The formula used is as follows:

[0084] ;

[0085] In the formula, This represents the adaptive pooling function. This represents a function that performs global average pooling down to a resolution of c×c. represents the learnable weight matrix, and x represents the input independent variable;

[0086] Step S322: Design an improved MobileNetV3 model, specifically by replacing the ReLU activation function with a differentiable pruning activation function in the feature calculation of each layer, and combining it with adaptive pooling. The formula used for feature calculation in each layer is as follows:

[0087] ;

[0088] In the formula, Indicates the first The output features of the layer Represents a differentiable shearing activation function. This represents the batch normalization function. Indicates the first -1 layer output features;

[0089] Step S323: Output layer design, specifically, using depthwise separable convolution combined with pointwise convolution to obtain visible light branch output features, the formula used is as follows:

[0090] ;

[0091] In the formula, This represents the output characteristics of the visible light branch. This represents depthwise separable convolution. This represents pointwise convolution. Indicates the first The output features of the layer This indicates the total number of layers calculated based on visible light branching features;

[0092] Step S33: Motion branch construction, used to capture the dynamic behavior patterns of people in the elevator scene. Specifically, it calculates the motion vectors of adjacent frames using the total variational-L1 optical flow algorithm, inputs them into the gated cyclic unit to model the temporal dependency, and obtains the aggregated motion features of 5 consecutive frames. The formula used is as follows:

[0093] ;

[0094] In the formula, Represents the motion vector at time t. Represents the total variational -L1 optical flow function. This represents the corrected data at time t. Indicates the characteristics of aggregation motion. This represents a gated loop unit function. This represents the motion vector at time t-2. This represents the motion vector at time t+2;

[0095] By performing the above operations, this solution addresses the technical problems of traditional intelligent elevator personnel recognition methods, which rely solely on single-frame image analysis or simple inter-frame difference, resulting in insufficient recognition capabilities for short-lived actions, continuous behaviors, and motion-blurred scenes. Furthermore, these methods are prone to feature distortion due to image blurring when personnel move quickly, making it difficult to construct a coherent behavioral logic chain and ultimately leading to an increased misjudgment rate in dynamic scenes. This solution creatively adopts a method of capturing continuous inter-frame motion features through motion branching and temporal modeling. It utilizes optical flow algorithms to quantize motion trajectories and combines recurrent networks to model long-term temporal dependencies, effectively distinguishing between real personnel actions and environmental interference, and enhancing stability for fast-moving and blurred frames.

[0096] Step S34: Construction of the graph convolution module, used to integrate visible light branch output features and aggregated motion features, and to construct a human joint relationship model that conforms to biomechanical principles based on graph convolution. Specifically, visible light branch output features and aggregated motion features are fused through dynamic gating weights, and the positions of major human joints are located through key point heatmaps. After graph convolution propagation, human perception features are obtained. The steps include:

[0097] Step S341: Dynamic gated feature fusion, specifically, compresses the feature channel dimension through global average pooling, generates dynamic gate weights using a two-layer fully connected network, and then weights and fuses the visible light branch output features and aggregated motion features. The formula used is as follows:

[0098] ;

[0099] In the formula, Gw represents the dynamic gating weight. This represents the sigmoid function. Represents the ReLU activation function. This represents the weights of the first fully connected layer. This represents the weights of the second fully connected layer. Represents the fusion feature, and Res represents the residual term, which is the element-wise sum of the visible light branch output feature and the aggregated motion feature. This indicates element-wise multiplication;

[0100] Step S342: Generating a key point heatmap. The key points correspond to human body key points in COCO format. The formula used to generate the key point heatmap is as follows:

[0101] ;

[0102] In the formula, Represents a heatmap of key points. This represents a convolution function with a kernel size of 1×1;

[0103] Step S343: Joint relationship modeling, specifically, modeling the nonlinear relationships between joints using a learnable adjacency matrix, with the following formula:

[0104] ;

[0105] In the formula, Represents the elements of the adjacency matrix. This represents the LeakyReLU activation function. This represents the learnable attention weight matrix. This represents the heatmap of the m-th key point. This represents the heatmap of the nth key point. This represents the heatmap of the k-th key point;

[0106] Step S344: Graph convolution propagation, using the following formula:

[0107] ;

[0108] In the formula, Indicates the first The layer graph convolution outputs features, where A represents the adjacency matrix. Indicates the first Layer graph convolution outputs features. Indicates the first Layer graph convolution weight matrix;

[0109] Step S35: Construct a multi-dimensional output module for synchronously outputting personnel behavior recognition, personnel quantity, and occlusion status. The steps include:

[0110] Step S351: Output of personnel behavior recognition, specifically, the personnel behavior recognition result is obtained through a fully connected layer based on the softmax function, using the following formula:

[0111] ;

[0112] In the formula, This indicates the results of personnel behavior recognition. This represents the softmax function. This represents the weight matrix output by personnel behavior recognition. This represents the bias term in the output of personnel behavior recognition. Indicates human sensory characteristics;

[0113] Step S352: Output the number of people. Specifically, this involves linearly mapping the keypoint heatmap after Softmax normalization to a density map to obtain the number of people. The formula used is as follows:

[0114] ;

[0115] In the formula, Den represents the density map. Represents the density mapping matrix, This indicates the number of people. This indicates the rounding operation. This represents the summation of all elements in the density plot;

[0116] Step S353: Occlusion state output, specifically, by performing max pooling on the human perception features and then classifying them using a multilayer perceptron, the occlusion state result is obtained. The formula used is as follows:

[0117] ;

[0118] In the formula, This indicates the result of the occlusion state. Represents the function of a multilayer perceptron;

[0119] Step S36: Construct and train the model. Specifically, construct a dual-branch deep network model by constructing the depth correction module, the visible light branch, the motion branch, the graph convolution module, and the multidimensional output module. Train the model based on the recognition training set and verify the model performance based on the recognition test set to obtain the dual-branch deep network model as the person recognition model.

[0120] By performing the above operations, this solution addresses the technical problems of traditional intelligent elevator personnel recognition methods, which struggle to capture complex posture changes and dynamic interactive behaviors, are susceptible to changes in lighting, occlusion interference, and viewpoint distortion, and are unable to model temporal action patterns and spatial joint relationships, resulting in high false detection rates in scenarios with dense crowds, rapid movement, or partial occlusion. This solution creatively adopts a dual-branch deep network model as the personnel recognition model. By fusing static features and dynamic temporal information, and combining multimodal data to model human behavior and spatial distribution in complex scenarios, it effectively solves the problem of decreased recognition accuracy caused by viewpoint changes, personnel occlusion, and dynamic blurring in elevator environments.

[0121] Example 5, see Figure 1 This embodiment is based on the above embodiment. In step S4, the elevator personnel identification specifically involves using the dataset to be identified as the input of the personnel identification model to perform intelligent elevator personnel identification, obtaining elevator personnel identification reference results, and comprehensively evaluating the personnel situation in the elevator based on the elevator personnel identification reference results. The elevator personnel identification reference results specifically include personnel behavior identification results, personnel quantity results, and occlusion status results.

[0122] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0123] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.

[0124] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A deep learning-based intelligent elevator personnel recognition method, characterized in that: The method includes the following steps: Step S1: Data collection inside the elevator. By collecting raw data, a raw dataset for personnel identification is obtained. The raw dataset for personnel identification specifically includes historical raw datasets and current raw datasets. Step S2: Data preprocessing. Through data preprocessing, we obtain the dataset to be identified, the recognition training set, and the recognition test set. Step S3: Personnel recognition model construction, used to construct the model required for personnel recognition in intelligent elevators, specifically to construct a dual-branch deep network model as the personnel recognition model. The dual-branch deep network model specifically includes a depth correction module, a visible light branch, a motion branch, a graph convolution module, and a multi-dimensional output module. The motion branch is used to capture the dynamic behavior patterns of people in the elevator scene. Specifically, it calculates the motion vectors of adjacent frames using the total variational-L1 optical flow algorithm, inputs them into the gated cyclic unit to model the temporal dependency, and obtains the aggregated motion features of 5 consecutive frames. The formula used is as follows: ; In the formula, Represents the motion vector at time t. Represents the total variational -L1 optical flow function. This represents the corrected data at time t. Indicates the characteristics of aggregation motion. This represents a gated loop unit function. This represents the motion vector at time t-2. This represents the motion vector at time t+2; Step S4: Elevator personnel identification, specifically, using the personnel identification model to identify elevator personnel and obtain reference results for personnel identification inside the elevator.

2. The intelligent elevator personnel recognition method based on deep learning according to claim 1, characterized in that: The construction of the personnel identification model specifically includes the following steps: Step S31: Construction of the depth correction module, used to further eliminate image distortion caused by installation angle and lens characteristics. Specifically, it involves constructing a multi-scale feature pyramid to analyze the image distortion pattern, generating a pixel-level deformation correction field, and obtaining a geometrically corrected image that conforms to the real spatial distribution. The steps include: Step S311: Multi-scale distortion analysis is used to capture the distortion features of different scales generated by the camera in the elevator scene. Specifically, the image pyramid features are extracted step by step through a five-layer coding structure, and the receptive field is expanded by dilated convolution to obtain the initial deformation field. Step S312: Deformation field regression, used to map abstract features into interpretable geometric correction parameters, specifically by fusing shallow details and deep semantic features through deconvolutional layers and skip connections to obtain an interpretable offset field; Step S313: Adaptive remapping, used to achieve high-precision image geometric correction, specifically by calculating the weighted pixel values ​​of the non-integer coordinate neighborhood in the input data, and combining the Gaussian kernel function to suppress interpolation noise, to obtain the correction result; Step S32: Visible light branch construction, used to capture the static features of people in the elevator scene. Specifically, it improves the depthwise separable convolutional structure of MobileNetV3, reduces computational cost while using a differentiable pruning activation function, and combines adaptive pooling to statistically determine channel importance. The steps include: Step S321: Design adaptive pooling to focus on the effective human body region in a noisy elevator scene. Specifically, channel importance is statistically analyzed using global average pooling, and channel attention weights are generated using Softmax. The formula used is as follows: ; In the formula, This represents the adaptive pooling function. This represents a function that performs global average pooling down to a resolution of c×c. represents the learnable weight matrix, and x represents the input independent variable; Step S322: Design an improved MobileNetV3 model, specifically by replacing the ReLU activation function with a differentiable pruning activation function in the feature calculation of each layer, and combining it with adaptive pooling; Step S323: Output layer design, specifically, using depthwise separable convolution combined with pointwise convolution to obtain visible light branch output features; Step S33: Construction of motion branches; Step S34: Construct the graph convolution module; Step S35: Construct a multi-dimensional output module for synchronously outputting personnel behavior recognition, personnel quantity, and occlusion status. The steps include: Step S351: Output of personnel behavior recognition, specifically by obtaining the personnel behavior recognition result through a fully connected layer based on the softmax function; Step S352: Output the number of people. Specifically, the key point heatmap is normalized by Softmax and then linearly mapped to a density map to obtain the number of people. Step S353: Occlusion state output, specifically by performing max pooling on human perception features and then classifying them using a multilayer perceptron to obtain the occlusion state result; Step S36: Construct and train the model. Specifically, construct a dual-branch deep network model by constructing the depth correction module, the visible light branch, the motion branch, the graph convolution module, and the multidimensional output module. Train the model based on the recognition training set and verify the model performance based on the recognition test set to obtain the dual-branch deep network model as the person recognition model.

3. The intelligent elevator personnel recognition method based on deep learning according to claim 2, characterized in that: The graph convolution module is constructed to integrate visible light branch output features and aggregated motion features, and to build a human joint relationship model that conforms to biomechanical principles based on graph convolution. Specifically, it fuses visible light branch output features and aggregated motion features through dynamic gating weights, locates the positions of major human joints through key point heatmaps, and obtains human perception features after graph convolution propagation. The steps include: Step S341: Dynamic gating feature fusion, specifically, compressing the feature channel dimension through global average pooling, generating dynamic gating weights using a two-layer fully connected network, and weighted fusing the visible light branch output features and aggregated motion features; Step S342: Generating a key point heatmap. The key points correspond to human body key points in COCO format. The formula used to generate the key point heatmap is as follows: ; In the formula, Represents a heatmap of key points. This represents a convolution function with a kernel size of 1×1. Indicates fusion characteristics; Step S343: Joint relationship modeling, specifically, modeling the nonlinear relationships between joints using a learnable adjacency matrix, with the following formula: ; In the formula, Represents the elements of the adjacency matrix. This represents the LeakyReLU activation function. This represents the learnable attention weight matrix. This represents the heatmap of the m-th key point. This represents the heatmap of the nth key point. This represents the heatmap of the k-th key point; Step S344: Graph convolution propagation.

4. The intelligent elevator personnel recognition method based on deep learning according to claim 1, characterized in that: Both the historical recognition raw dataset and the current recognition raw dataset include key point coordinate data collected by the depth camera and elevator monitoring video data. The historical recognition raw dataset also includes personnel posture annotation data, personnel number annotation data, and personnel occlusion annotation data.

5. The intelligent elevator personnel recognition method based on deep learning according to claim 1, characterized in that: The data preprocessing is used to preprocess the collected raw data, and specifically includes the following steps: Step S21: Video data processing, used to eliminate video image distortion, ensure video temporal consistency and focus video key frames, specifically to perform geometric distortion correction, time synchronization alignment and key frame extraction on the elevator monitoring video data; Step S22: Personnel posture data processing, used to process the joint coordinate data collected by the depth camera. Specifically, manual quality inspection is used to correct the labeling error of the joint coordinate data collected by the depth camera, and the bone length is normalized according to the national standard. Step S23: Optical flow data optimization, used to extract the precise motion vectors of people in the elevator scene. Specifically, it restores the image blur caused by the rapid movement of people by using a blurred Gaussian kernel function and fast Fourier transform, and filters out noise by median filtering and direction consistency detection. Step S24: Dataset segmentation, used to segment the historical dataset into a training set and a test set, specifically, to segment the original historical recognition dataset into a recognition training set and a recognition test set; Step S25: Perform preprocessing, specifically by preprocessing the current recognition raw dataset through video data processing, personnel posture data processing, and optical flow data optimization to obtain the dataset to be recognized; and by preprocessing the historical recognition raw dataset through video data processing, personnel posture data processing, optical flow data optimization, and dataset segmentation to obtain the recognition training set and the recognition test set.

6. The intelligent elevator personnel recognition method based on deep learning according to claim 1, characterized in that: The elevator personnel identification specifically involves using the dataset to be identified as input to the personnel identification model to perform intelligent elevator personnel identification, obtaining elevator personnel identification reference results, and comprehensively evaluating the personnel situation inside the elevator based on the elevator personnel identification reference results. The elevator personnel identification reference results specifically include personnel behavior identification results, personnel quantity results, and occlusion status results.