Passenger intention recognition method and system based on elevator scene-behavior association
By applying a passenger intention recognition method based on elevator scene-behavior association in the elevator system, using I3D network and space-time visual Transformer, the problem that traditional elevator management systems cannot respond to passenger needs in real time is solved, achieving higher recognition accuracy and response efficiency.
Patent Information
- Application Number
- CN202510143105.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Traditional elevator management systems are unable to respond to passengers' dynamic needs in real time and lack accurate identification of passengers' intentions, which leads to long-term waiting and safety hazards during peak hours.
Through a passenger intention recognition method based on elevator scene-behavior association, using an I3D network and a spatiotemporal visual Transformer, the passenger's motion trajectory and interaction characteristics are captured and modeled, and context information is injected to enhance the association between the scene and the behavior subject.
It improves the accuracy and response efficiency of passenger intention recognition, enhances the intelligence level of the elevator system, can more accurately understand passengers' riding needs, and achieve personalized response.
Smart Images

Figure CN120220010A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for passenger intention recognition based on elevator scenario-behavior association, belonging to the field of computer vision. Background Art
[0002] Traditional elevator management and scheduling rely on fixed algorithms and cannot respond to the dynamic needs of passengers in real time. Especially during peak hours, passengers often face long waiting times, leading to congestion and potential safety hazards. In addition, traditional systems lack accurate recognition of passenger intentions and have relatively insufficient response capabilities in emergency situations. Therefore, improving the intelligence level of elevator systems has become an important task that needs to be solved urgently.
[0003] In recent years, the rapid development of technologies such as deep learning and computer vision has made the research and application of intelligent elevators increasingly concerned. The advantages of deep learning in video behavior recognition can help elevator systems identify the behaviors and intentions of passengers, thereby optimizing scheduling strategies and improving operation efficiency. At the same time, as a key computer vision technology, it provides a safer and more convenient service for elevators by real-time monitoring of passenger behaviors.
[0004] In the fields of computer vision and artificial intelligence, video behavior recognition refers to the process of analyzing and reasoning a given video sequence to obtain its category label. At the same time, it is noted that there are intricate explicit or implicit relationship patterns in video data. The essence of behavior modeling is exactly how to understand the various relationships existing within the video. However, existing methods still have many deficiencies in understanding and modeling these relationships. The method of extracting the features of the entire scene depends to a large extent on the appearance features of video frames, which is prone to introducing inductive biases. The method of extracting and independently modeling the appearance and spatial position changes of active objects to understand actions is not affected by object or appearance biases, but lacks the additional clues provided by the context for the interaction between objects. These methods ignore the association between key information such as the scene and the behavior subject in scene relationship understanding, resulting in difficult to achieve consistent performance in actual complex scenes (such as intelligent elevator scenes), bringing certain challenges to the research. Summary of the Invention
[0005] In order to enhance the association between the scene and the behavior subject, make the model pay more attention to interactive objects, improve the accuracy and robustness of the behavior recognition model, especially in the intelligent elevator lobby scene, by analyzing the behaviors of passengers, their intentions can be deeply understood, so as to more accurately grasp the elevator-taking needs of passengers, achieve personalized response of the elevator system, and improve its response efficiency. The present invention provides a method and system for passenger intention recognition based on elevator scenario-behavior association, and the technical solutions are as follows:
[0006] The passenger intention recognition method based on elevator scenario-behavior association of the present invention includes:
[0007] Step 1: Obtain video frames, detect targets in the video, and generate a detection box representation of the object;
[0008] Step 2: Use the I3D network to extract local features of the detection box, obtain the feature map of the target detection box area as the region of interest, so as to obtain the position information and appearance features of each pedestrian in each frame;
[0009] Step 3: Concatenate the local features obtained in Step 2 with the position features obtained by mapping the detection box coordinates to obtain the aggregated features of each object in each frame;
[0010] Step 4: Introduce a motion trajectory aggregation module to integrate the aggregated features extracted for each object in the video sequence across time, so as to capture and construct the motion trajectory of each object over time, and reason about different motion trajectories along the time dimension to obtain a motion trajectory representation with time features, and model the interaction features between them based on the motion trajectories of different objects;
[0011] Step 5: Use the interaction token obtained in Step 4 as the query of the context injection module, use the original video frame as the key and value, input it into the spatio-temporal vision Transformer, and update the trajectory token with the information contained in the video token through the cross-attention mechanism.
[0012] Optionally, Step 2 includes:
[0013] Step 21: Use the I3D network to extract spatial and temporal information to obtain the feature map of the video frame;
[0014] Step 22: According to the detection box coordinates (x, y, w, h), convert them into ROI coordinates corresponding to the I3D feature map, and the upper left and lower right coordinates are expressed as:
[0015] (x1, y1) = (x - w / 2, y - h / 2)
[0016] (x2, y2) = (x + w / 2, y + h / 2)
[0017] Among them, x and y are the center point coordinates of the detection box, and w and h are the width and height of the detection box;
[0018] Step 23: Perform bilinear interpolation on each sub-region (i, j) of the ROI, which is expressed as:
[0019]
[0020] Among them, F is the resolution of the I3D feature map, wk and w l is the weight for the sub-region;
[0021] Step 24: Pool the eigenvalues of all sub-regions within the ROI region into a feature vector of a fixed size through average pooling.
[0022] Optionally, the motion trajectory with temporal features obtained in step 4 is represented as:
[0023]
[0024] z i = MLP(h i )
[0025] where h i represents the motion trajectory of the i-th object, and z i represents the trajectory token of the i-th object.
[0026] Optionally, step 5 includes:
[0027] Step 51: Through the spatial cross-attention mechanism, compare the interaction query q t and the context key k st , and find the best matching position of the interaction trajectory in space, represented as:
[0028]
[0029] where s represents the spatial position, t represents the time frame, q t = z i W q , k st = x st W k , v st = x st W v , W q , W k , W v represent weights;
[0030] Step 52: In the time dimension, use the self-attention mechanism to process the trajectory tokens, capture the interaction features in a single frame, fuse multi-frame information, and form a complete trajectory representation with a global view.
[0031] The present invention provides a passenger intention recognition system based on elevator scenario-behavior association, and the system includes:
[0032] A target detection module, configured to obtain video frames, detect targets in the video, and generate a detection box representation of the object;
[0033] The local feature extraction module is configured to use the I3D network to perform local feature extraction on the detection box, obtain the feature map of the target detection box area as the region of interest, so as to obtain the position information and appearance features of each pedestrian in each frame;
[0034] The feature aggregation module is configured to splice the local features obtained by the local feature extraction module and the position features obtained by mapping the detection box coordinates to obtain the aggregated features of each object in each frame;
[0035] The motion trajectory aggregation module is configured to perform cross-time integration on the aggregated features extracted for each object in the video sequence to capture and construct the motion trajectory of each object over time, and perform reasoning on different motion trajectories along the time dimension to obtain a motion trajectory representation with time features, and model the interaction features between them based on the motion trajectories of different objects;
[0036] The context injection module is configured to use the interaction token obtained by the motion trajectory aggregation module as the query of the context injection module, and use the original video frame as the key and value, input it into the spatio-temporal vision Transformer, and update the trajectory token with the information contained in the video token through the cross-attention mechanism.
[0037] Optionally, the processing process of the local feature extraction module includes:
[0038] First, use the I3D network to extract spatial and temporal information to obtain the feature map of the video frame;
[0039] Secondly, according to the detection box coordinates (x, y, w, h), convert them into ROI coordinates corresponding to the I3D feature map, and the upper left and lower right coordinates are expressed as:
[0040] (x1, y1) = (x - w / 2, y - h / 2)
[0041] (x2, y2) = (x + w / 2, y + h / 2)
[0042] Where x and y are the center point coordinates of the detection box, and w and h are the width and height of the detection box;
[0043] Then, perform bilinear interpolation on each sub-region (i, j) of the ROI, expressed as:
[0044]
[0045] Where F is the resolution of the I3D feature map, w k and w l are the weights for the sub-region;
[0046] Finally, the eigenvalue of all sub-regions within the ROI region is aggregated into a feature vector of a fixed size through average pooling.
[0047] Optionally, the motion trajectory with temporal features obtained by the motion trajectory aggregation module is expressed as:
[0048]
[0049] z i = MLP(h i )
[0050] where h i represents the motion trajectory of the i-th object, and z i represents the trajectory token of the i-th object.
[0051] Optionally, the processing process of the context injection module includes:
[0052] By means of the spatial cross-attention mechanism, compare the interactive query q t and the context key k st , and find the best matching position of the interactive trajectory in space, expressed as:
[0053]
[0054] where s represents the spatial position, t represents the time frame, q t = z i W q , k st = x st W k , v st = x st W v , W q , W k , W v represent weights;
[0055] In the time dimension, the self-attention mechanism is used to process the trajectory tokens, capture the interactive features in a single frame, fuse the information of multiple frames, and form a complete trajectory representation with a global view.
[0056] The present invention provides a passenger intention recognition device based on elevator scenario-behavior association, including a memory and a processor;
[0057] The memory is used to store computer programs;
[0058] The processor is used to implement the method described in any one of the above when executing the computer program.
[0059] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the above is implemented.
[0060] The beneficial effects of the present invention are:
[0061] The present invention proposes a method for decoupling object motion and background context, which is used to inject context into captured object interactions. A target detection module is adopted to represent the motion trajectories of different objects in different frames in the form of detection boxes, capture the motion trajectories of pedestrians to model interactions, understand the moving subjects and their interactions with each other, avoid the deviation caused by the appearance changes of people and objects due to action execution, and enable the model to better focus on pedestrians who may have the intention of taking the elevator; a method for context injection is proposed, using interaction tokens as queries of the vision Transformer to inject context information into the interaction representation, thereby enhancing the association between the scene and the behavioral subjects. The interaction trajectories after context injection can more accurately capture the complex interactions between objects, thus improving the understanding of the scene and behavioral intentions, and providing more accurate decision-making support for applications such as intelligent elevator systems.
[0062] Experimental results show that compared with existing methods, the present invention effectively improves the accuracy of scene and behavioral intention recognition. Description of the Drawings
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0064] Figure 1 It is a network framework diagram of the intelligent elevator passenger intention recognition method based on video surveillance of the present invention.
[0065] Figure 2 It is a visualization diagram of pedestrian target detection in the intelligent elevator scenario of the present invention.
[0066] Figure 3 It is a visualization diagram of pedestrian intention recognition in the intelligent elevator scenario of the present invention. Detailed Embodiments
[0067] To make the purpose, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the drawings.
[0068] Embodiment 1:
[0069] This embodiment provides a method for recognizing passenger intentions based on elevator scenario-behavior association. Refer to Figure 1 , the method includes:
[0070] Step 1: Detect the target and generate a detection box representation of the object, which is used to represent the pedestrians in the detected video and their position coordinates.
[0071] This embodiment uses YOLOv10 to detect each object in the video frame and obtain the coordinate information of its detection box.
[0072] YOLOv10 achieves state-of-the-art performance while significantly reducing computational overhead and meeting the excellent accuracy-latency trade-off by eliminating non-maximum suppression (NMS) and optimizing various model components from the perspectives of efficiency and accuracy. It mainly includes the following steps:
[0073] Step 11: Read the input video file, extract frames from the video at a frame rate of 6 to obtain a set of consecutive video frames with a size of H×W.
[0074] Step 12: Divide the input image into S×S grids, where S is the size of the network output feature map. When the input image is 640×640 pixels and the downsampling ratio of the network is 32, then that is, the image is divided into 20×20 grids, and the size of each grid unit is:
[0075]
[0076] On the image, these grids cover the entire image area. When the center point of an object falls into a certain grid unit, then this unit will be responsible for predicting the bounding box position, size, confidence, and category of the object.
[0077] Step 13: Use the enhanced CSPNet as the backbone network to extract image features. This network extracts deep features of the image through multiple convolutional layers and residual connections, which can reduce gradient flow and computational redundancy to improve the network's feature extraction ability and efficiency.
[0078] Step 14: Input the obtained features into the connection layer to integrate features at different scales for detecting targets of different sizes and pass them to the output part of the network. It internally integrates the PAN (Path Aggregation Network) layer to achieve effective fusion of multi-scale features.
[0079] Meanwhile, YOLOv10 has two prediction heads. One is a one-to-many prediction head that generates multiple predictions for each object during training to provide rich supervision signals for improving the accuracy of learning. It does not take effect during the inference stage, thus reducing the computational load. The other is a one-to-one prediction head that generates one optimal prediction for each object during inference, eliminating the need for the NMS (Non-Maximum Suppression) operation, thereby reducing latency and improving inference efficiency. The model uses both prediction heads during training, enabling it to utilize the rich supervision signals of one-to-many assignment during training, while using the prediction results of one-to-one assignment during inference, thus achieving efficient inference without NMS. To achieve prediction-aware matching for the two branches, a consistent matching metric is utilized. By adjusting the matching metric parameters, the supervision signals of one-to-one and one-to-many assignments are made consistent, reducing the supervision gap during training and enhancing the prediction quality of the model.
[0080] Step 15: The detection box for each object obtained contains the center point coordinates (x, y) of the detection box, width and height (w, h), confidence p, and the probability C for each category. i 。
[0081] Step 2: Extract local features.
[0082] Obtain the feature map of the target detection box area as the region of interest, thereby obtaining the position information and appearance features of each pedestrian in each frame.
[0083] In this embodiment, the I3D network is used to extract features from video frames, and ROI Align is adopted to map the target detection boxes generated by YOLOv10 onto the I3D feature map for extracting the features of the corresponding regions. First, the I3D network is used to extract spatial and temporal information to obtain the feature map of the video frame. Second, according to the bounding box coordinates (x, y, w, h) output by YOLOv10, they are converted into ROI coordinates corresponding to the I3D feature map, and the upper left and lower right coordinates can be expressed as:
[0084] (x1, y1) = (x - w / 2, y - h / 2)
[0085] (x2, y2) = (x + w / 2, y + h / 2)
[0086] Among them, x and y are the center point coordinates of the detection box, and w and h are the width and height of the detection box.
[0087] ROI Align uses bilinear interpolation method to accurately obtain the feature values within the ROI from the feature map, rather than simply using integer coordinates. This step ensures the spatial smoothness and accuracy of the features. The ROI region can be represented as (x1, y1, x2, y2), which is divided into an M×N grid. Bilinear interpolation is performed on each sub-region (i, j), which can be expressed as:
[0088]
[0089] where F is the resolution of the I3D feature map, w k and w l are the weights for the sub-region, usually the reciprocal of the distance:
[0090]
[0091] Then, the feature values of all sub-regions within the ROI region are pooled into a fixed-size feature vector through average pooling.
[0092] Step 3: Concatenate the local features obtained in Step 2 with the position features obtained by mapping the detection box coordinates to obtain the aggregated features of each object in each frame.
[0093] Step 4: Introduce a motion trajectory aggregation module to perform cross-time integration on the aggregated features extracted for each object in the video sequence, in order to capture and construct the motion trajectories of each object over time, and reason about different motion trajectories along the time dimension to understand their interaction relationships.
[0094] The motion trajectory aggregation module is an important part of temporal modeling for object interaction, especially when understanding and analyzing the interaction behaviors between objects. Although the local feature extraction module can obtain the aggregated features of objects, these features are often static and lack continuity in the time dimension, thus limiting the in-depth understanding of the dynamic changes in object behaviors. To make up for this deficiency, it is necessary to perform cross-time integration on the aggregated features extracted for each object in each frame of the video sequence, in order to capture and construct the motion trajectories of each object over time. These trajectory features not only contain the appearance features of the objects, but also cover their spatial position information and evolution over time, and can be expressed as:
[0095]
[0096] where X represents the motion trajectory features of N objects, represents the local aggregated feature of the Nth object in the Tth frame.
[0097] In each frame, the aggregated feature x of a given object is provided, and over time, the features of each object are further integrated to gain a deeper understanding of the spatio-temporal dynamics of the objects in the video. This method utilizes the positional encoding of the aggregated features to directly connect the features of the same object at different time points, thereby obtaining a motion trajectory representation with temporal features:
[0098]
[0099] z i = MLP(h i )
[0100] where h i represents the motion trajectory of the i-th object, and z i represents the trajectory token of the i-th object.
[0101] Step 5: Inject the original video frame as context information into the interaction features to enhance the association between the scene and the actors.
[0102] In the absence of background visual context, traditional object motion trajectory modeling can usually only capture the basic correlations between people and objects. However, context information plays a crucial role in capturing interactions between objects because these interactions are often not affected by the appearance biases of single objects or the environment. By introducing context information, additional clues can be provided for object interactions, helping to more accurately infer the intentions and behaviors of people. For example, walking towards an elevator may imply the next action of wanting to take the elevator. Therefore, injecting context information into the context-separated trajectory tokens, refining the trajectory-centered interaction representation with information from the background, and injecting important information from interactions between people and objects, objects and objects, thereby improving the understanding of behaviors.
[0103] This embodiment designs a context injection module to enrich the feature representation of the trajectory tokens by introducing a spatio-temporal visual Transformer. Specifically, the trajectory token z is used as a query and input into the spatio-temporal visual Transformer together with the video token x st (as keys and values), and the trajectory tokens are updated with the information contained in the video tokens through the cross-attention mechanism.
[0104] First, through the spatial cross-attention mechanism, comparing the interaction query q t and the context key k st , the best matching position of the interaction trajectory in space is found, which can be expressed as:
[0105]
[0106] where s represents the spatial position, t represents the time frame, and q t = zi W q ,k st = x st W k ,v st = x st W v ,W q ,W k ,W v represents the weight.
[0107] Next, aggregate the interaction trajectories across the time dimension to infer the interconnections between different time steps. In the time dimension, use the self-attention mechanism to process the trajectory tokens to further refine their interaction information. This process enables the trajectory tokens to not only capture the interaction features in a single frame but also span the entire video sequence, fuse the information of multiple frames, and form a complete trajectory representation with a global view. Finally, the interaction trajectories after injecting context can more accurately capture the complex interactions between objects, thereby enhancing the understanding of the scene and behavioral intentions and providing more accurate decision-making support for applications such as intelligent elevator systems.
[0108] Embodiment 2
[0109] This embodiment provides a method for recognizing passenger intentions based on elevator scene-behavior association. As Figure 1 shown, a set of RGB video frames in an elevator scene is used as the input to the model.
[0110] First, use the YOLOv10 network to perform object detection on pedestrians in the elevator scene to obtain the position coordinates and class labels of different pedestrians in each frame. Among them, the video is defaultly framed at a frame rate of 6, and each frame is adjusted to a resolution of 224×224. The results of object detection are in Figure 1It is represented in the figure that it includes the object detection box, class label, and probability score. Subsequently, the local feature extraction module uses the I3D network initialized after pre-training on the Kinetics-400 dataset as the backbone for feature extraction, and applies ROI Align to extract the features of each individual in the given N bounding box regions in each video frame feature map. The individual features are concatenated with the position features obtained by mapping the detection box coordinates to obtain the aggregated features of each object in each frame. Next, using the motion trajectory aggregation module, the static aggregated features of each object in each frame of the video sequence output by the local feature extraction module are integrated along the time dimension to capture and construct the motion trajectory of each object over time, and model the interactions between them based on the motion trajectories of different objects. Finally, the obtained interaction tokens are used as the query of the context injection module, and the original video frames are used as the key and value, and are input into the spatio-temporal vision Transformer together to achieve the injection of context information. By matching the trajectory token query with the context video token key, the spatial attention mechanism first calculates the best position of the interaction trajectory. Next, the temporal attention mechanism performs cross-time interaction trajectory pooling to accumulate temporal information in the interaction tokens. Then the interaction tokens with context injection are used for action recognition.
[0111] The experimental results of this embodiment are all verified on the publicly available Something-Else dataset and the self-built elevator scene behavior intention dataset. The Something-Else dataset is an extension of the Something-Something-V2 dataset and is designed for compositional action recognition. Compositional action recognition aims to decompose each human action into a combination of one or more verbs, subjects, and objects, ensuring that the action elements between the training and test sets do not overlap, and emphasizing their independence and composability. It also aims to understand the relationship between them by separating the interaction between humans and objects from their background and appearance deviations. By achieving this goal, the machine can obtain insights that help better generalize to new environments.
[0112] The dataset contains 174 action categories and 112,795 videos, divided into 54,919 for training and 57,876 for validation, all using a combined setting. In this task, there are two disjoint sets of nouns (objects) {A, B} and two disjoint sets of verbs (actions) {1, 2}. During training, the model can observe combinations of nouns and verbs from one set, while during testing, different combinations are used. Specifically, during training, the model can observe objects from {1A + 2B}, while during testing, objects from {1B + 2A} are used. This setting aims to identify new verb-noun combinations during testing. Performance evaluation follows a standard classification setting, which includes metrics such as top-1 and top-5 accuracy.
[0113] The self-built dataset of elevator scene behavior intentions aims to focus on the recognition of pedestrian behavior intentions in real elevator environments, providing in-depth insights into human interactions in elevator scenarios. This dataset focuses on various interactions of pedestrians in front of and around elevators in different elevator scenarios, covering a variety of possible behaviors and intentions. Specifically, it includes four main action categories, namely: walking towards the elevator and taking the elevator, passing by the elevator, staying near the elevator but not taking it, and leaving the elevator. These categories not only reflect the common behavior patterns of people in elevator environments but also help understand their behavior intentions in specific situations. This dataset contains a total of 1558 video samples, divided into 1067 for training and 491 for validation. All action categories are covered in both the training set and the validation set to ensure the comprehensiveness and effectiveness of model training. In addition, the videos record the elevator usage at different times and levels of congestion, thus ensuring that the dataset can cover various possible real-world scenarios. These scenarios include busy elevator environments during peak hours and relatively static usage situations, reflecting the behavioral differences of people in different situations. Performance evaluation follows a standard classification setting, which includes metrics such as top-1 and top-3 accuracy.
[0114] Adopting the technology provided by the present invention has the following characteristics:
[0115] (1) The present invention proposes a method of decoupling motion and context to model human interactions. By capturing the motion trajectories of different objects, it models the relationships between people and objects, and between objects and objects, avoiding the deviation caused by the appearance changes of people and objects due to action execution.
[0116] (2) The present invention proposes a method of context injection, using interaction tokens as queries for the vision Transformer to inject context information into the interaction representation, thereby enhancing the association between the scene and the behavioral entity.
[0117] (3) Experimental results on the Something-Else dataset show the effectiveness of the present invention in the combined action recognition task. It was also evaluated on a self-built dataset of elevator scene behavior intentions, and significant performance improvements were achieved compared to the baseline model.
[0118] To verify the effectiveness of different modules for action recognition, this embodiment was first verified on the Something-Else dataset, and the results are shown in Tables 1 and 2.
[0119] Table 1 Comparison of model component performance on the Something-Else dataset
[0120]
[0121] Table 2 Comparison of model component performance on the elevator scene behavior intention dataset
[0122]
[0123] Motionformer was selected as the context injection module and initialized with pre-trained weights on the Something-SomethingV2 dataset, and STIN was selected as the motion trajectory aggregation module. It can be observed that compared with using the Motionformer model alone for behavior recognition, the accuracy of the present invention has been significantly improved, achieving top-1 and top-5 accuracies of 8% and 4.9% respectively. Compared with using the STIN+I3D model alone for behavior recognition, the top-1 and top-5 accuracies have been improved by 13% and 10.1% respectively, easily surpassing the baseline.
[0124] In Figure 2 shows the visualization results of object detection in the elevator scene. YOLOv10 was used as the object detection module and initialized with pre-trained weights on the COCO dataset. It can be observed that the present invention can detect the pedestrian category and its location (represented in the form of a detection box) with high accuracy even from a wide-range and long-distance perspective, and can associate the position changes of the same person across frames to generate corresponding motion trajectories. It provides accurate and rich prior information for predicting whether pedestrians outside the elevator car have a need to take the elevator.
[0125] In Figure 3The visualization results of intent recognition in the elevator scenario are shown. It can be migrated to the intelligent elevator scenario with only minor adjustments to predict the elevator-riding intent of pedestrians in the elevator scenario. By utilizing the pedestrian movement trajectories and combining additional clues from the context, the interactions and relationships among different pedestrians are observed, and the final prediction of the elevator-riding intent of pedestrians outside the elevator car (represented by intent labels and probability scores) is obtained, so as to facilitate the implementation of personalized responses of the elevator system and improve the response efficiency of the elevator system.
[0126] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.
[0127] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A passenger intention recognition method based on elevator scene-behavior association, characterized in that: The method comprises: Step 1: Get the video frame, detect the target in the video, and generate the detection box representation of the object; Step 2: Use the I3D network to extract local features of the detection frame, obtain the feature map of the target detection frame area as the region of interest, and obtain the position information and appearance features of each pedestrian in each frame; Step 3: Concatenate the local features obtained in step 2 with the position features obtained by mapping the detection frame coordinates to obtain the aggregated features of each object in each frame; Step 4: Introduce the motion trajectory aggregation module to integrate the aggregated features extracted from each object in the video sequence across time to capture and construct the motion trajectory of each object over time, and infer different motion trajectories along the time dimension to obtain a motion trajectory representation with time characteristics, and model the interaction features between different objects based on their motion trajectories; Step 5: The interaction token obtained in step 4 is used as a query for the context injection module, and the original video frame is used as the key and value, input into the spatiotemporal visual Transformer, and the trajectory token is updated with the information contained in the video token through the cross-attention mechanism.
2. The method according to claim 1, characterized in that The step 2 comprises: Step 21: Use the I3D network to extract spatial and temporal information to obtain a feature map of the video frame; Step 22: According to the detection frame coordinates (x, y, w, h), convert them into ROI coordinates corresponding to the I3D feature map, and the coordinates of the upper left corner and the lower right corner are expressed as: (x1,y1)=(xw / 2,yh / 2) (x2,y2)=(x+w / 2,y+h / 2) Among them, x, y are the coordinates of the center point of the detection box, w, h are the width and height of the detection box; Step 23: Perform bilinear interpolation on each sub-region (i, j) of the ROI, expressed as: Where F is the resolution of the I3D feature map, w k and w l is the weight for the sub-region; Step 24: The feature values of all sub-regions in the ROI region are aggregated into a feature vector of a fixed size through average pooling.
3. The method according to claim 1, characterized in that: The motion trajectory with time characteristics obtained in step 4 is expressed as: With i =MLP(h i ) Among them, h i represents the motion trajectory of the i-th object, z i The trajectory token representing the i-th object.
4. The method according to claim 1, characterized in that: The step 5 comprises: Step 51: Compare the interactive query q through the spatial cross attention mechanism t and context key k st , find the best matching position of the interaction trajectory in space, expressed as: Among them, s represents the spatial position, t represents the time frame, and q t =z i W q , k st =x st W k , v st =x st W v , W q , W k , W v represents weight; Step 52: In the time dimension, the self-attention mechanism is used to process the trajectory tokens, capture the interactive features in a single frame, fuse multi-frame information, and form a complete trajectory representation with a global perspective.
5. A passenger intention recognition system based on elevator scene-behavior association, characterized in that: The system comprises: The object detection module is configured to obtain a video frame, detect an object in the video, and generate a detection box representation of the object; A local feature extraction module is configured to use an I3D network to perform local feature extraction on the detection frame, obtain a feature map of the target detection frame area as the region of interest, and thereby obtain position information and appearance features of each pedestrian in each frame; A feature aggregation module is configured to concatenate the local features obtained by the local feature extraction module with the position features obtained by mapping the detection frame coordinates to obtain aggregated features of each object in each frame; The motion trajectory aggregation module is configured to integrate the aggregated features extracted from each object in the video sequence across time to capture and construct the motion trajectory of each object over time, and to infer different motion trajectories along the time dimension to obtain a motion trajectory representation with time characteristics, and to model the interaction characteristics between different objects based on their motion trajectories; The context injection module is configured to use the interaction token obtained by the motion trajectory aggregation module as a query of the context injection module, and the original video frame as a key and value, input into the spatiotemporal visual Transformer, and update the trajectory token with the information contained in the video token through the cross-attention mechanism.
6. The system according to claim 5, characterized in that The processing process of the local feature extraction module includes: First, the I3D network is used to extract spatial and temporal information to obtain the feature map of the video frame; Secondly, according to the detection frame coordinates (x, y, w, h), they are converted into ROI coordinates corresponding to the I3D feature map, and the coordinates of the upper left corner and the lower right corner are expressed as: (x1,y1)=(xw / 2,yh / 2) (x2,y2)=(x+w / 2,y+h / 2) Among them, x, y are the coordinates of the center point of the detection box, w, h are the width and height of the detection box; Then, bilinear interpolation is performed on each sub-region (i, j) of the ROI, expressed as: Where F is the resolution of the I3D feature map, w k and w l is the weight for the sub-region; Finally, the feature values of all sub-regions in the ROI area are aggregated into a feature vector of fixed size through average pooling.
7. The system according to claim 5, characterized in that The motion trajectory with time characteristics obtained by the motion trajectory aggregation module is expressed as: With i =MLP(h i ) Among them, h i represents the motion trajectory of the i-th object, z i The trajectory token representing the i-th object.
8. The system according to claim 5, characterized in that The processing process of the context injection module includes: Comparing interactive queries q by spatial cross attention mechanism t and context key k st , find the best matching position of the interaction trajectory in space, expressed as: Among them, s represents the spatial position, t represents the time frame, and q t =z i W q , k st =x st W k , v st =x st W v , W q , W k , W v represents weight; In the temporal dimension, a self-attention mechanism is used to process trajectory tokens, capture interactive features in a single frame, and fuse multi-frame information to form a complete trajectory representation with a global perspective.
9. A passenger intention recognition device based on elevator scene-behavior association, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is configured to implement the method according to any one of claims 1 to 4 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Video classification method and system for intelligent elevator passenger intention analysis
CN118155119A