Elevator detection maintenance operation supervision method and system based on spatio-temporal behavior recognition
By using multi-view video stream data processing and behavioral semantic graph matching technology, the problems of low efficiency and insufficient accuracy in elevator inspection and maintenance supervision have been solved, realizing full-process supervision and automated compliance judgment of elevator inspection and maintenance operations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SPECIAL EQUIP SAFETY SUPERVISION INSPECTION INST OF JIANGSU PROVINCE
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-03
AI Technical Summary
Existing elevator inspection and maintenance supervision methods rely on manual on-site spot checks, which are inefficient, highly subjective, and difficult to cover the entire process. Furthermore, they lack a deep semantic understanding of the behavior of operators, the use of tools, key operating actions, and the order of operations, making it impossible to accurately identify violations.
By using multi-view video stream data processing, combined with target detection and attitude estimation models, spatiotemporal behavioral features of operators and tools are extracted, a semantic graph of operational behavior is constructed, and matched with standard operation templates to generate compliance judgment results and risk levels.
It enables continuous monitoring of the entire elevator inspection and maintenance process, improves the accuracy and stability of key operation identification, and can automatically determine problems such as omissions, incorrect sequences, and misuse of tools, generating standardized behavior labels and risk warning information.
Smart Images

Figure CN122336434A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of elevator maintenance technology, and in particular to a method and system for supervising elevator inspection and maintenance operations based on spatiotemporal behavior recognition. Background Technology
[0002] As frequently used special equipment, elevators' inspection and maintenance quality directly affects operational and public safety. Current elevator inspection and maintenance supervision methods mainly rely on manual on-site spot checks, verification of paper records, and post-operation video review, which is still largely a reactive supervision model. This approach is insufficient to cover the entire inspection and maintenance process and is highly dependent on the experience of supervisors, resulting in low efficiency, strong subjectivity, and difficulty in timely detection of violations.
[0003] Existing technologies attempt to record elevator operation and maintenance processes through equipment sensors, IoT data acquisition terminals, or ordinary video surveillance. However, these solutions typically only obtain elevator operating status, environmental information, or simple video footage, lacking the ability to deeply understand the behavior of operators, tool usage, key operational actions, and work sequences. Therefore, existing technologies struggle to achieve stable, accurate, and automated identification of typical violations such as missed inspections, incorrect operation sequences, tool misuse, inadequate key actions, and non-standard postures.
[0004] Furthermore, elevator inspection and maintenance operations are characterized by high continuity, fine-grained movements, dispersed work objects, frequent obstructions, and significant scene variations. Traditional methods based on single-frame target detection or simple video recognition often only recognize coarse-grained movements, making it difficult to accurately understand the interaction relationships between "personnel—tools—components—areas—states," and failing to form a complete chain of evidence for accountability and review. Therefore, how to construct an intelligent supervision method and system that can address the entire elevator inspection and maintenance process, taking into account behavior recognition, semantic understanding of key movements, compliance judgment, and risk warning, has become a pressing technical problem to be solved in this field. Summary of the Invention
[0005] In view of this, the purpose of this invention is to propose an elevator inspection and maintenance operation supervision method and system based on spatiotemporal behavior recognition that can take into account behavior recognition, key action semantic understanding, compliance judgment and / or risk warning in the process of elevator maintenance work supervision.
[0006] To achieve the above-mentioned technical objectives, the technical solution adopted by this invention is as follows: A method for supervising elevator inspection and maintenance operations based on spatiotemporal behavior recognition, comprising: S01. Acquire multi-view video streams of the elevator inspection and maintenance work area, and collect work task data, elevator operation status data and equipment component identification data corresponding to the multi-view video streams. Then perform spatiotemporal alignment processing on them to obtain work input data. S02. Based on the target detection model, the video frames in the work input data are processed to identify feature targets including workers, and the human body key point information of workers is extracted based on the posture estimation model. The feature target trajectory is constructed by combining target association tracking. S03. Based on the feature target trajectory, the video stream is segmented into candidate action segments, and the candidate action segments are input into the I3D model and the SlowFast model to extract static spatial features and dynamic temporal features to obtain a spatiotemporal behavior feature vector. S04. Based on the inter-frame difference results and the change results of action recognition confidence, locate key frames for candidate action segments, and extract key action semantic descriptors by combining semantic segmentation results, human body key point information and target contact relationship. S05. Construct a job behavior semantic graph based on the spatiotemporal behavior feature vector, key action semantic descriptors and temporal sequence relationships; S06. Call the standard operation template corresponding to the current task, match the operation behavior semantic graph with the standard operation template to obtain the operation judgment result, and generate a compliance judgment result and risk level based on the operation judgment result.
[0007] As one possible implementation, further, in this solution S01, the elevator inspection and maintenance work area includes the car area, landing door area, machine room area and / or pit area.
[0008] As a possible implementation, further, in this solution S01, the multi-view video stream, the operation task data, the elevator operation status data, and the equipment component identification data are time-synchronized and operation task-bound to obtain spatiotemporally aligned operation input data; The spatiotemporally aligned task input data also includes ladder number information, task type information, and component location code information corresponding to the current task.
[0009] As one possible implementation, further, in this solution S02, the characteristic targets include worker targets, work tool targets, and elevator operating component targets; The characteristic target trajectory includes the trajectory of the operator, the trajectory of the operating tool, and the interaction trajectory of the operating components.
[0010] As one possible implementation, further, in this solution S02, the method for identifying worker targets, work tool targets, and elevator operating component targets based on the target detection model includes: The YOLOv8 target detection model is used to detect workers, wrenches, measuring tools, protective equipment and elevator components in video frames. The cross-frame target correspondence is established through a multi-target association tracking algorithm to generate target trajectory sequences.
[0011] As a possible implementation, further, in this solution S02, the method for extracting the human body key point information of the worker based on the posture estimation model includes: The OpenPose posture estimation model is used to extract key points of the head, shoulders, elbows, wrists, hips, knees and / or ankles of the worker. Combined with the relative position changes of key points, joint angle changes and key point velocity changes, posture state features used to characterize the work posture are obtained.
[0012] As a possible implementation, further, in this solution S04, the key action semantic descriptor includes at least action category information, tool category information, active component information, contact state information, human posture state information and / or action duration information; The contact state information is determined by the spatial proximity and overlap relationships between the tool target area, the key point area of the human hand, and the operating component area.
[0013] As a possible implementation, further, in this solution S05, the job behavior semantic graph at least represents the association between the operator, the job tool, the operating component, the scene area, the action event, and the equipment status.
[0014] As one possible implementation, further, in this solution S05, the nodes in the job behavior semantic graph include personnel nodes, tool nodes, component nodes, area nodes, action nodes and / or status nodes; The edges in the semantic graph of the job behavior include edges related to temporal sequence, contact, spatial adjacency, action dependency, and / or state constraint.
[0015] As a possible implementation, further, in this solution S06, the standard operation template is pre-set with one or more of the following: a set of prescribed actions for the corresponding operation item, a prescribed sequence of actions, a correspondence between tools and components, a range of allowed durations for actions, human body safety posture constraints, and elevator operating status constraints. The task judgment results include one or more of the following: task continuity result, action sequence result, tool usage consistency result, posture standardization result, and action completion result; The compliance determination result includes at least one of the following: normal, missing items, out of order, misuse of tools, abnormal posture, and incomplete key actions.
[0016] As a preferred implementation method, this solution further includes: S07. Based on the compliance assessment results and risk level, output standardized behavior labels, risk warning information and / or traceability evidence data.
[0017] As a preferred implementation method, the traceability evidence data in this solution preferably includes keyframe images, key action timestamps, standardized behavior tag sequences, compliance scoring results, risk event records, and / or index information corresponding to the original video clips.
[0018] Based on the above, this solution also proposes an elevator inspection and maintenance operation monitoring system based on spatiotemporal behavior recognition, which includes: The multi-source data acquisition module is used to collect multi-view video streams, work task data, elevator operating status data, and equipment component identification data from the elevator inspection and maintenance work area. The data preprocessing module is used to perform time synchronization, task binding, and candidate action segmentation on the multi-view video stream, task data, elevator operation status data, and equipment component identification data. The target parsing module is used to identify targets such as workers, tools, and elevator operating components, and to extract key information about the workers' bodies. The spatiotemporal behavior recognition module is used to extract static spatial features and dynamic temporal features from candidate action segments based on the I3D model and the SlowFast model, and obtain spatiotemporal behavior feature vectors. The key action extraction module is used to extract key action semantic descriptors of candidate action segments based on inter-frame difference results, action recognition confidence change results, semantic segmentation results, human key point information and target contact relationship. The semantic graph construction module is used to construct a job behavior semantic graph based on the spatiotemporal behavior feature vector, key action semantic descriptors, and temporal sequence relationships. The assessment and early warning module is used to match the semantic graph of the work behavior with the standard operation template and output the compliance judgment result and risk level. The data archiving and output module is used to output standardized behavioral labels, risk warning information, and traceability evidence data.
[0019] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art: This solution unifies the spatiotemporal alignment of multi-view video streams with work task data and elevator operation status data, expanding the monitored objects from discrete images to work events with task context. This enables it to cover multiple work areas such as the car, landing doors, machine room, and pit, achieving continuous monitoring of the entire elevator inspection and maintenance process and significantly enhancing the overall monitoring capability.
[0020] This solution combines YOLOv8 target detection, OpenPose pose estimation, I3D spatiotemporal modeling, and SlowFast dual-rate temporal modeling. It can extract the spatial relationships between personnel, tools, and components, and characterize the rhythm, duration, and state changes of actions. Therefore, it can more accurately distinguish fine-grained continuous actions in elevator maintenance, improve the accuracy and stability of key operation action identification, and significantly enhance the ability to identify complex fine-grained actions in the data.
[0021] This solution also goes beyond simply classifying isolated actions by using keyframe localization, key action semantic extraction, and the construction of a semantic graph of work behavior. Instead, it further expresses the work sequence, tool correspondence, posture standardization, action completion degree, and equipment state constraints, thereby enabling automated judgment of issues such as omissions, out-of-order items, misuse of tools, abnormal posture, and incomplete key actions. This elevates the compliance judgment of this solution from identifying actions to understanding the process.
[0022] This solution can output standardized behavior tags, risk levels, and violation information during the operation, and simultaneously generate keyframe evidence, timestamp indexes, and original video-related data to form a complete chain of evidence. Compared with manual review, this solution can significantly improve the efficiency of risk discovery, the ability to trace responsibility, and the ability to close the regulatory loop. It has strong engineering practical value. Overall, this solution has strong real-time early warning and traceability verification capabilities in the supervision of elevator inspection and maintenance operations. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a simplified implementation flowchart of the regulatory methods in this plan; Figure 2 This is a schematic diagram of the unit module connections of the monitoring system in this solution. Detailed Implementation
[0025] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] like Figure 1 As shown in the figure, this embodiment of the method for monitoring elevator inspection and maintenance operations based on spatiotemporal behavior recognition includes: S01. Acquire multi-view video streams of the elevator inspection and maintenance work area, and collect work task data, elevator operation status data and equipment component identification data corresponding to the multi-view video streams. Then perform spatiotemporal alignment processing on them to obtain work input data. S02. Based on the target detection model, the video frames in the work input data are processed to identify feature targets including workers, and the human body key point information of workers is extracted based on the posture estimation model. The feature target trajectory is constructed by combining target association tracking. S03. Based on the feature target trajectory, the video stream is segmented into candidate action segments, and the candidate action segments are input into the I3D model and the SlowFast model to extract static spatial features and dynamic temporal features to obtain a spatiotemporal behavior feature vector. S04. Based on the inter-frame difference results and the change results of action recognition confidence, locate key frames for candidate action segments, and extract key action semantic descriptors by combining semantic segmentation results, human body key point information and target contact relationship. S05. Construct a job behavior semantic graph based on the spatiotemporal behavior feature vector, key action semantic descriptors and temporal sequence relationships; S06. Call the standard operation template corresponding to the current task, match the operation behavior semantic graph with the standard operation template to obtain the operation judgment result, and generate a compliance judgment result and risk level based on the operation judgment result.
[0027] As one possible implementation, further, in this solution S01, the elevator inspection and maintenance work area includes the car area, landing door area, machine room area and / or pit area.
[0028] In this solution S01, the multi-view video stream, the work task data, the elevator operation status data, and the equipment component identification data are synchronized in time and bound to the work task to obtain spatiotemporally aligned work input data; wherein, the spatiotemporally aligned work input data also includes elevator number information, work item type information, and component location code information corresponding to the current work task.
[0029] Elevator inspection and maintenance work typically occurs in multiple dispersed areas such as the machine room, car, landing doors, and pit, and the same task often spans multiple operational scenarios. In existing technologies, video surveillance data, work order data, and elevator operating status data are fragmented, causing the monitoring system to only "see the screen" but unable to accurately determine whether the screen belongs to a specific maintenance task, and making it difficult to associate subsequently identified actions with specific elevator numbers, components, or work orders. Therefore, the first step is to address the alignment issue of multi-source heterogeneous data within a unified monitoring timeline and task context.
[0030] As an example, step S01 of this solution may include the following: Multiple camera terminals are pre-deployed at the elevator inspection and maintenance site, covering at least the machine room area, car area, landing door area, and pit area. Each camera terminal continuously collects video streams and uploads video frames to edge computing nodes or a monitoring platform. Simultaneously, the system collects the following auxiliary data from the maintenance management terminal, elevator control interface, and work order management system: (1) Current work order number; (2) Current ladder number identifier; (3) Current task type; (4) Elevator operating status; (5) Current elevator component identification information; (6) The start time and the planned end time of the task; (7) Scene area number corresponding to the camera terminal.
[0031] In order to enable multiple video streams, work order information, and elevator status information to enter the same regulatory semantic space, this solution first synchronizes the time of each video source, and then binds the synchronized video segments with the current work order execution time period, elevator number, and scene area to obtain a unified spatiotemporally aligned operation input dataset.
[0032] In terms of time alignment of multiple video streams, the first... The first camera terminal collected the first Frame image is Its local timestamp is Because each camera terminal has a clock skew, it is necessary to estimate the time offset of each video stream relative to the reference clock. .
[0033] In order to estimate First, calculate the motion energy sequence of each video stream. For the ... Road video, defining the first The motion energy of a frame is:
[0034] in, Indicates the first Road video in the first The motion energy of a frame; Indicates the effective pixel area of the image; Indicates the first Road Video No. Frame in pixel coordinates The grayscale value or brightness value at that location.
[0035] Then, using the reference camera terminal Using this as a baseline, the time offset is calculated using the principle of maximizing cross-correlation:
[0036] in: Indicates the first The time offset of the video path relative to the reference video; Indicates the candidate offset; The reference video is in the [number]th [section]. Motion energy of a frame.
[0037] Therefore, the first The local timestamp of the video feed has been corrected to a unified timestamp:
[0038] in: Indicates the first timeline under a unified timeline Road Video No. The timestamp of the frame.
[0039] The above reasoning is based on the fact that the same maintenance action will cause similar motion energy peak sequences from different perspectives. Therefore, the cross-perspective time deviation can be obtained by cross-correlation of the maximum position.
[0040] Based on the above, work orders are bound to achieve data spatiotemporal alignment. Let the current work order to be bound be... Its task time range is The corresponding ladder number for the operation is Allow work area set as Then for the first... Road Video No. Frame, defining its relationship with work order The binding determination function is:
[0041] in: Indicates the first Road Video No. Does the frame belong to a work order? ; This represents an indicator function, which takes the value 1 if the condition within the parentheses is true, and 0 otherwise. Indicates the first The elevator number corresponding to the video feed; Indicates the first The scene area number corresponding to the video feed.
[0042] when At that time, it indicates that the video frame is related to the work order. Since they are consistent in terms of time, tier number, and region, they should be included in the current regulatory tasks.
[0043] In this step, after time synchronization and task binding, the spatiotemporal aligned job input dataset is obtained, which is defined as follows:
[0044] in: This represents the input dataset for the spatiotemporal alignment task; This indicates the current work order identifier, which is used to subsequently call the corresponding standard operation template; Represents timestamp The corresponding elevator operation status data is used to subsequently determine whether the timing of a certain action is compliant; timestamp Used for subsequent trajectory continuity analysis and motion duration calculation; Indicates the first The elevator number corresponding to the video feed; Indicates the first The scene area number corresponding to the video feed is used to determine the location where the action occurred. For the first The first camera terminal collected the first The frame image is used for subsequent object detection and pose estimation.
[0045] As one possible implementation, further, in this solution S02, the characteristic targets include worker targets, work tool targets, and elevator operating component targets; The characteristic target trajectory includes the trajectory of the operator, the trajectory of the operating tool, and the interaction trajectory of the operating components.
[0046] In this solution S02, the method for identifying worker targets, work tool targets, and elevator operating component targets based on the target detection model includes: The YOLOv8 target detection model is used to detect workers, wrenches, measuring tools, protective equipment and elevator components in video frames. The cross-frame target correspondence is established through a multi-target association tracking algorithm to generate target trajectory sequences.
[0047] In this scheme S02, the method for extracting human key point information of workers based on the posture estimation model includes: The OpenPose posture estimation model is used to extract key points of the head, shoulders, elbows, wrists, hips, knees and / or ankles of the worker. Combined with the relative position changes of key points, joint angle changes and key point velocity changes, posture state features used to characterize the work posture are obtained.
[0048] Step S01 of this solution binds the video to the work order, but it is still not possible to directly understand "who is working, what tools they are using, and which parts they are operating" in the video. Since people, tools, and parts in elevator maintenance scenarios are frequently partially occluded, have changing perspectives, and switch positions, it is necessary to first establish a stable target detection, posture analysis, and trajectory association mechanism to solve the problem of computable representation of the work subject, work medium, and work object.
[0049] Step S02 of this solution will generate the dataset output in step S01. Each frame of the image is input into the YOLOv8 object detection network to identify the bounding boxes of workers, tools, and elevator components. Then, the target region of each worker is input into the OpenPose human pose estimation network to extract key points such as the head, shoulders, elbows, wrists, hips, knees, and ankles.
[0050] To ensure continuous recognition across frames, it is necessary to perform correlation matching between adjacent frames for the same person, tool, and component to establish a trajectory sequence. Simultaneously, to determine whether the "person-tool-component" interaction constitutes a valid operation, it is also necessary to calculate key hand points, the proximity between the tool and component, and construct an interaction relationship sequence.
[0051] As an example, step S02 of this solution may include the following: 1. Representation of target detection results For the The target set is obtained by analyzing the frame images using YOLOv8:
[0052] in: Indicates the first The set of target detection results for the frame; Indicates the first The total number of targets detected in the frame; Indicates the first The bounding box of each target; Indicates the first Category labels for each target, which may include workers, wrenches, measuring instruments, protective equipment, elevator parts, etc. Indicates the first The detection confidence of each target.
[0053] 2. Representation of key points in human posture For the The first worker, in the first Extract the set of human key points from the frame:
[0054] in: Indicates the first The first worker was on the A set of key points of the human body in a frame; Indicates the total number of key points; Indicates the first The coordinates of the key points; Indicates the first Detection confidence of key points.
[0055] 3. Trajectory Association Model To establish cross-frame trajectories, a correlation cost matrix is constructed by combining existing trajectories from the previous frame with the detection results of the current frame. For the trajectory... With current detection target The cost is defined as:
[0056] in: Representing the trajectory With current detection target The associated costs; This represents the intersection-union ratio (IUU) between the predicted bounding box and the detected bounding box. Representing the trajectory At the predicted position in the current frame; Indicates the first The bounding box of each target; Representing the trajectory The appearance feature vector; Indicates the current detection target The appearance feature vector; Representing the trajectory The motion state vector; Indicates the current detection target The motion observation vector; Indicates the weights of each item, and .
[0057] Since the same target should simultaneously satisfy the requirements of similar spatial location, similar appearance features, and continuous motion trend in adjacent frames, modeling these three factors together can improve trajectory stability.
[0058] By optimally allocating the cost matrix, the following trajectory set is established:
[0059]
[0060]
[0061] in: Represents the set of worker trajectories; Represents a set of tool trajectories; Represents the set of trajectories of the operating components; , , These represent the time-series trajectories of a specific person, tool, or component.
[0062] 4. Human-Tool-Component Interaction Scoring Model Since it is necessary to determine whether the task action actually occurred, an interaction score is defined. Let the first... The workers were at the The center of the left and right hand key points of the frame is ,tool The center point is ,part The center point is The interaction score is then defined as:
[0063] in: Indicates the first People in the frame ,tool and components Interaction score; This represents the Euclidean distance between the hand and the tool. Indicates the Euclidean distance between the tool and the component; , This represents the distance attenuation coefficient.
[0064] when A larger value indicates that the person's hand is close to the tool, and the tool is close to the component, suggesting a higher likelihood of actual operation.
[0065] In actual maintenance videos, only a portion of the video contains truly valuable operational segments, with a large number of frames representing waiting, walking, turning, and static background phases. Directly performing fine-grained behavior recognition on the entire video would result in computational redundancy and easily introduce invalid noise. Therefore, step S03 of this solution accurately segments candidate action segments from continuous video and extracts spatiotemporal features from these segments that distinguish fine-grained operational actions.
[0066] Step S03 of this scheme, based on the personnel trajectory, tool trajectory, component trajectory, and interaction score sequence obtained in step S02, first calculates the motion activity level at each moment. For time intervals where the activity level is continuously higher than the threshold, these intervals are divided into candidate motion segments. Then, each candidate motion segment is fed into the I3D network and the SlowFast network respectively to extract spatiotemporal features that take into account spatial appearance, motion rhythm, and temporal evolution. Finally, the posture statistical features and interaction statistical features obtained in step S02 are concatenated into the depth features to form a unified spatiotemporal behavior feature vector.
[0067] As an example, step S03 of this solution may include the following: 1. Establish a candidate action activity model For the Frames define the activity level of an action. for:
[0068] in: Indicates the first Frame motion activity; Indicates the trajectory of the worker in the th position. The velocity vector of the frame; Indicates the speed modulus of personnel movement; Indicates the tool trajectory at the 1st The velocity vector of the frame; Indicates the speed modulus of tool movement; Indicates the first Average interaction score of the human-tool-component frame; Indicates the first Frame relative to the first The amount of change in human joint posture in a frame; Represents the weighting coefficients, and .
[0069] in, It can be obtained through the changes in the included angles of several key joints. Let the coordinates of the three points, shoulder, elbow, and wrist, be... The joint angle is defined as follows:
[0070] Further calculation of the average change at multiple key joints yields... .
[0071] when continuous Frames greater than the threshold At that time, the corresponding time interval is defined as a candidate action segment:
[0072] in: Indicates the first One candidate action segment; For the first Frame image; Indicates the starting frame of this segment; This indicates the end frame of the segment.
[0073] Since real-world operations are usually accompanied by personnel displacement, tool movement, enhanced interaction, and posture changes, combining these four types of quantities with weights can provide a relatively stable way to locate high-value action segments.
[0074] 2. Dual-path feature extraction using I3D and SlowFast methods For candidate action segments First, the data is fed into the I3D network to obtain spatial-short-term dynamic fusion features:
[0075] in: Representing fragments Feature vectors extracted by the I3D network; This represents the I3D feature extraction mapping.
[0076] Then, construct two sets of inputs for the same segment using slow sampling and fast sampling respectively, and feed them into the SlowFast network to obtain:
[0077]
[0078] in: This represents a slow-sampled sequence of candidate action segments, primarily preserving scene structure and semantics at longer time scales; This represents a fast sampling sequence of candidate action segments, primarily preserving the rhythm and short-term variations of the action. This indicates the output characteristics of the slow branch; This indicates the output characteristics of fast branching.
[0079] 3. Fusion of spatiotemporal behavioral feature vectors The depth features are concatenated with the interaction statistical features and pose statistical features output from step S02 to obtain the final spatiotemporal behavior feature vector:
[0080] in: Indicates the first The final spatiotemporal behavior feature vector of each candidate action segment; Represents the feature fusion weight matrix; Represents the bias vector; This represents a vector concatenation operation; This represents the statistical features of human pose within the candidate fragment; This represents the statistical features of interaction scores within candidate segments.
[0081] Further, perform initial action classification on this feature:
[0082]
[0083] in: Represents the log-probability vector for action category classification; Represents the classification weight matrix; Represents the classification bias vector; This represents the probability distribution of candidate action segments belonging to each action category.
[0084] As a possible implementation, further, in this solution S04, the key action semantic descriptor includes at least action category information, tool category information, active component information, contact state information, human posture state information and / or action duration information; The contact state information is determined by the spatial proximity and overlap relationships between the tool target area, the key point area of the human hand, and the operating component area.
[0085] Although step S03 has identified candidate action segments, the focus of elevator maintenance supervision is not on the average characteristics of the entire segment, but on the key moments within the segment that truly embody "start of operation, tool contact, and completion of action." Existing technologies, if only performing segment-level classification, are prone to overlooking crucial evidence and are insufficient to support subsequent compliance determinations. Therefore, step S04 precisely extracts key frames with judgment value and their semantic information from the candidate action segments.
[0086] In step S04, for each candidate action segment, the inter-frame difference intensity, action category confidence change intensity, and interaction state change intensity are calculated and fused to form a keyframe scoring sequence. Based on the scoring sequence, the action start keyframe, action peak keyframe, and action end keyframe are located.
[0087] After keyframe localization is completed, a key action semantic descriptor is generated by combining the tool type, the type of the active component, the contact state, the human posture state, and the duration of the action. This descriptor can serve as direct input for the subsequent construction of the semantic graph of the task behavior.
[0088] As an example, step S04 of this solution may include the following: 1. Inter-frame differential intensity For the candidate action segment, the first Frame, the inter-frame differential strength is defined as:
[0089] in: Indicates the first Inter-frame differential intensity; Indicates the first The effective analysis area of a frame is preferably the area jointly surrounded by personnel, tools, and components; Indicates the first Frame in pixel coordinates The pixel value at that location.
[0090] 2. Intensity of change in action confidence Since classification confidence often changes significantly when the action actually occurs, the following definition is given:
[0091] in: Indicates the first Frame and the The intensity of action confidence changes between frames; Represents the calculation of the first based on the sliding window. Frame action probability distribution; This represents the L1 norm.
[0092] 3. Intensity of interactive changes To depict the essential nature of the operation—the hand touching the tool and the tool touching its components—we define the intensity of the interaction change:
[0093] in: Indicates the first Frame and the Inter-frame interaction score change; This represents the average interaction score for the nth frame.
[0094] 4. Keyframe scoring function After normalizing the above three quantities, a weighted fusion is performed to obtain the keyframe score:
[0095] in: Indicates the first Keyframe scoring; These represent the normalized inter-frame difference intensity, action confidence change intensity, and interaction change intensity, respectively. Represents the weighting coefficients, and .
[0096] The reasoning logic for the scoring function in this step is as follows: When the action just begins, the inter-frame difference and interaction changes will increase significantly; When the core action occurs, the action confidence and interaction intensity usually reach their peak. When the action ends, all quantities will return to normal.
[0097] Therefore, joint scoring can be used to robustly locate the start, peak, and end keyframes of an action.
[0098] Let the first The keyframes for each action segment are:
[0099]
[0100]
[0101] in: Indicates the first The starting keyframe of an action segment; Indicates the first Peak keyframes of a single action segment; Indicates the first The end keyframe of an action segment; , This indicates the threshold for keyframe determination.
[0102] 5. Basic quantities for posture standardization To construct subsequent semantic descriptors, the human pose state needs to be extracted from keyframes. Taking the shoulder-elbow-wrist points as an example, the joint angles are:
[0103] in: Indicated by point The joint angle at the vertex; This represents the coordinates of the corresponding key points. By statistically analyzing multiple joint angles of the upper limbs, trunk, and lower limbs, the posture state feature vector can be obtained. .
[0104] 6. Key Action Semantic Descriptors Construct the first [action type] by combining action type, tool type, active part, contact state, posture state, and duration. Semantic descriptors for each action segment:
[0105] in: Indicates the first Key action semantic descriptors for each action segment; Indicates the action category; Indicates the tool category; Indicates the category of functional components; Indicates the contact state; Indicates posture state characteristics; Indicates the duration of the action.
[0106] The compliance of elevator maintenance operations does not depend on whether a single action is identified, but rather on whether multiple actions meet the prescribed execution sequence, tool correspondence, component correspondence, and state constraint relationships. Isolated action labels alone cannot support full-process supervision. Therefore, step S05 of this solution organizes the discrete detection results into a structured semantic graph that can express process relationships and constraint relationships.
[0107] As a possible implementation, further, in this solution S05, the job behavior semantic graph at least represents the association between the operator, the job tool, the operating component, the scene area, the action event, and the equipment status.
[0108] As a possible implementation, further, in this solution S05, the nodes in the job behavior semantic graph include personnel nodes, tool nodes, component nodes, region nodes, action nodes and / or state nodes; the edges in the job behavior semantic graph include temporal sequence relationship edges, contact relationship edges, spatial adjacency relationship edges, action dependency relationship edges and / or state constraint relationship edges.
[0109] Step S05 of this scheme constructs a job behavior semantic graph based on the personnel, tool, and component trajectories obtained in step S02, and the spatiotemporal behavior feature vectors and key action semantic descriptors obtained in steps S03 and S04. This graph includes personnel nodes, tool nodes, component nodes, region nodes, action nodes, and state nodes, and establishes edge connections based on temporal sequence, contact relationships, spatial adjacency relationships, action dependencies, and state constraints.
[0110] Furthermore, step S05 of this solution also considers the work order type in step S01. (Work task data) Retrieve the standard behavior diagram template corresponding to the maintenance project from the standard work template library for matching in subsequent steps.
[0111] As an example, step S05 of this solution may include the following: Define the job behavior semantic graph as follows:
[0112] in: Semantic graph representing job behavior; Represents a set of nodes; Represents the set of edges; Represents a collection of node attributes; This represents the adjacency weight matrix.
[0113] 1. Construction of a node set
[0114] in: Represents a set of personnel nodes; Represents a set of tool nodes; Represents a set of component nodes; Represents a set of region nodes; Represents a set of action nodes; This represents the set of device status nodes.
[0115] For the Each action node has the following node attributes defined:
[0116] in: Indicates the first An attribute vector for each action node; This represents the spatiotemporal behavior feature vector obtained in step S3; This represents the key action semantic descriptor obtained in step S4; Indicates the start time of the action; Indicates the end time of the action.
[0117] 2. Edge weight construction For nodes and Its edge weight can be defined as:
[0118] in: Represents a node and The overall correlation weight; Indicates the weight of chronological order; Indicates the weight of contact relationships; Indicates the weight of spatial adjacency relationships; Indicates the weight of state constraint relationships; Represents the weighting coefficients, and .
[0119] The weight of the chronological order can be defined as follows:
[0120] in: , Indicates the time of the node event; This represents the time decay constant.
[0121] The contact relationship weights can be obtained by accumulating the interaction score sequence in step S02:
[0122] in: Indicates the number of frames included in the cumulative count; Represents a node The corresponding time range of the event; This represents the interaction score at that moment. This indicates a calculated index.
[0123] 3. Standard Operating Procedure Template For work orders For the corresponding maintenance project, the standard diagram is retrieved from the template library:
[0124] in: Indicates work order Corresponding standard operating procedure template; These represent the nodes, edges, node attributes, and adjacency weights of a standard graph, respectively.
[0125] After steps S01-S05, this solution has achieved the identification of actions, tools, components, postures, and process relationships. However, whether the elevator inspection and maintenance work complies with standards and specifications remains the core issue of this solution. In other words, the identification results need to be transformed into regulatory conclusions such as "missing items, incorrect sequence, tool misuse, abnormal posture, and incomplete key actions." Therefore, step S06 of this solution addresses the mapping issue from semantic recognition results to work compliance judgment results; step S06 of this solution transforms the semantic map of the actual work behavior obtained in step S05... Corresponding standard template diagram A step-by-step matching process is performed, calculating sub-scores based on five dimensions: action sequence, consistency of tool use, standardization of human posture, degree of action completion, and equipment status compatibility. These sub-scores are then weighted and combined to form a comprehensive compliance score. Finally, a risk level is generated based on the comprehensive compliance score and the specific type of violation.
[0126] As a possible implementation, further, in this solution S06, the standard operation template is pre-set with one or more of the following: a set of prescribed actions for the corresponding operation item, a prescribed sequence of actions, a correspondence between tools and components, a range of allowed durations for actions, human body safety posture constraints, and elevator operating status constraints. The task judgment results include one or more of the following: task continuity result, action sequence result, tool usage consistency result, posture standardization result, and action completion result; The compliance determination result includes at least one of the following: normal, missing items, out of order, misuse of tools, abnormal posture, and incomplete key actions.
[0127] As an example, step S06 of this solution may include the following: 1. Compliance of work sequence Assuming from the actual diagram The extracted action sequence is From the template diagram The standard action sequence extracted is The order conformity is defined as follows:
[0128] in: Indicates the degree of conformity of the work sequence; This represents the edit distance between the actual action sequence and the standard action sequence; Indicates the actual number of actions; Indicates the number of standard actions.
[0129] If there are missing, out-of-order, or redundant actions in the sub-steps, the editing distance will increase and the sequence compliance will decrease.
[0130] 2. Consistency in tool usage Let the first The tool-component pair corresponding to each actual action is: The corresponding standard pair in the template is Then, tool consistency is defined as:
[0131] in: The tool indicates that consistency scores are used; Indicates the type of tool actually used; Indicates the category of the actual functional component; Indicates the tool category and component category required by the template.
[0132] 3. Posture Correctness Score Let the first The number of joint angles that need to be examined in each movement is: , No. The actual angle of each joint is The center value of the reference angle interval specified in the template is The allowable deviation is The posture normativity score is defined as follows:
[0133] in: Indicates the first The score for the proper posture of each movement; Indicates the first Each joint allows for a normalized range scale to avoid the influence of different joint angles on the dimensionality.
[0134] Then, averaging over all actions, we get:
[0135] in: This indicates the score for the correct posture throughout the entire task.
[0136] 4. Score for completion of the action For the key actions required by the template, if the corresponding start keyframe, peak keyframe, and end keyframe are all detected, and the duration is within the allowed range, then... If the action is completed within a certain timeframe, then the action is considered finished. Definition:
[0137] in: Indicates the first A measure of the completion rate of an action; Represents the empty set; Indicates the duration of the actual action; This indicates the upper and lower limits of the action duration specified by the template.
[0138] Overall completion rate:
[0139] in: This indicates the score for the completion of the action.
[0140] 5. State Matching Score Some maintenance actions can only be performed in specific elevator states, such as inspection mode, elevator stop state, and power outage state. Therefore, let's assume the... When an action occurs, the system state is: The set of allowed states for the template is ,but:
[0141]
[0142] in: Indicates the first The status matching flag for each action; This represents the state matching score for the entire task.
[0143] 6. Overall Compliance Score The final overall compliance score is defined as follows:
[0144] in: This indicates the overall compliance score; Indicates the degree of conformity in the order; The tool indicates that consistency scores are used; This indicates the score for posture conformity. This indicates the score for the completion of the action; Indicates the state matching score; This represents the corresponding weight, and their sum is 1.
[0145] 7. Risk Level Calculation To further characterize the urgency of regulation, a risk score is defined:
[0146] in: Indicates risk score; This indicates the overall compliance score; This represents the severity coefficient of the violation, and is assigned a value based on the violation category, such as missing items, out-of-order items, tool misuse, abnormal posture, or failure to complete key actions. This represents the trigger coefficient for dangerous situations; a higher value is used if an illegal action is performed while the situation is dangerous. This indicates the risk score weight.
[0147] Then according to The system outputs low-risk, medium-risk, and high-risk levels based on a preset threshold.
[0148] As a preferred implementation method, this solution further includes: S07. Based on the compliance assessment results and risk level, output standardized behavior labels, risk warning information and / or traceability evidence data.
[0149] As a preferred implementation method, the traceability evidence data in this solution preferably includes keyframe images, key action timestamps, standardized behavior tag sequences, compliance scoring results, risk event records, and / or index information corresponding to the original video clips.
[0150] Numerical scores alone are insufficient to meet regulatory requirements. Typically, when reviewing violations, back-end supervisors need to see which specific action violated regulations, when it occurred, which keyframe it corresponds to, which elevator it belongs to, and which work order it pertains to. Therefore, this solution further proposes step S07, which organizes the aforementioned identification and judgment results into a searchable, alertable, and traceable chain of regulatory evidence.
[0151] Step S07 of this solution generates a standardized behavior tag sequence based on the comprehensive compliance score, risk score, and violation category output in step S06, and triggers an early warning message when the risk score exceeds a threshold. Simultaneously, keyframe images, key action timestamps, original video clip indexes, work order numbers, ladder numbers, and compliance score results are packaged into evidence records and saved to the regulatory platform database, forming a complete work archive.
[0152] As an example, step S07 of this solution may include the following: Let the first Evidence records of an action event are defined as follows:
[0153] in: Indicates the first Evidence of each action event; Indicates the elevator number; Indicates the work order number; Indicates standardized behavior labels; These represent the start keyframe time, peak keyframe time, and end keyframe time, respectively. This indicates the overall compliance score; Indicates risk score; This indicates the index address of the video segment corresponding to the action; This represents the summary checksum of the video clip.
[0154] The video clip summary checksum can be generated using a hash method:
[0155] in: Indicates the first Summary values of a video clip; This represents the video segment data corresponding to the action; This represents a hash operation.
[0156] The purpose of this step is to ensure that even if liability is subsequently traced, it can be based on... Quickly locate the original video, based on Verify whether the evidence fragment has been tampered with.
[0157] When the risk score exceeds the preset threshold When a warning event is triggered:
[0158] in: Indicates the first Does the action trigger an alert? This indicates the risk warning threshold.
[0159] This step, by outputting behavioral tags, scoring results, risk levels, keyframe evidence, and original video indexes, can form a complete chain of regulatory evidence. This enables the solution to not only automatically identify, but also automatically determine, automatically issue warnings, and automatically leave traces, meeting the dual needs of full-process supervision of elevator inspection and maintenance and post-event accountability and review.
[0160] Combination Figure 2 As shown, based on the above, this solution also proposes an elevator inspection and maintenance operation monitoring system based on spatiotemporal behavior recognition, which includes: The multi-source data acquisition module is used to collect multi-view video streams, work task data, elevator operating status data, and equipment component identification data from the elevator inspection and maintenance work area. The data preprocessing module is used to perform time synchronization, task binding, and candidate action segmentation on the multi-view video stream, task data, elevator operation status data, and equipment component identification data. The target parsing module is used to identify targets such as workers, tools, and elevator operating components, and to extract key information about the workers' bodies. The spatiotemporal behavior recognition module is used to extract static spatial features and dynamic temporal features from candidate action segments based on the I3D model and the SlowFast model, and obtain spatiotemporal behavior feature vectors. The key action extraction module is used to extract key action semantic descriptors of candidate action segments based on inter-frame difference results, action recognition confidence change results, semantic segmentation results, human key point information and target contact relationship. The semantic graph construction module is used to construct a job behavior semantic graph based on the spatiotemporal behavior feature vector, key action semantic descriptors, and temporal sequence relationships. The assessment and early warning module is used to match the semantic graph of the work behavior with the standard operation template and output the compliance judgment result and risk level. The data archiving and output module is used to output standardized behavioral labels, risk warning information, and traceability evidence data.
[0161] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0162] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0163] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for supervising elevator inspection and maintenance operations based on spatiotemporal behavior recognition, characterized in that, include: S01. Acquire multi-view video streams of the elevator inspection and maintenance work area, and collect work task data, elevator operation status data and equipment component identification data corresponding to the multi-view video streams. Then perform spatiotemporal alignment processing on them to obtain work input data. S02. Based on the target detection model, the video frames in the work input data are processed to identify feature targets including workers, and the human body key point information of workers is extracted based on the posture estimation model. The feature target trajectory is constructed by combining target association tracking. S03. Based on the feature target trajectory, the video stream is segmented into candidate action segments, and the candidate action segments are input into the I3D model and the SlowFast model to extract static spatial features and dynamic temporal features to obtain a spatiotemporal behavior feature vector. S04. Based on the inter-frame difference results and the change results of action recognition confidence, locate key frames for candidate action segments, and extract key action semantic descriptors by combining semantic segmentation results, human body key point information and target contact relationship. S05. Construct a job behavior semantic graph based on the spatiotemporal behavior feature vector, key action semantic descriptors and temporal sequence relationships; S06. Call the standard operation template corresponding to the current task, match the operation behavior semantic graph with the standard operation template to obtain the operation judgment result, and generate a compliance judgment result and risk level based on the operation judgment result.
2. The elevator inspection and maintenance operation supervision method based on spatiotemporal behavior recognition as described in claim 1, characterized in that, In S01, the elevator inspection and maintenance work area includes the car area, landing door area, machine room area and / or pit area; In S01, the multi-view video stream, the task data, the elevator operation status data, and the equipment component identification data are synchronized in time and bound to the task data to obtain spatiotemporally aligned task input data. The spatiotemporally aligned task input data also includes ladder number information, task type information, and component location code information corresponding to the current task.
3. The elevator inspection and maintenance operation supervision method based on spatiotemporal behavior recognition as described in claim 1, characterized in that, In S02, the characteristic targets include worker targets, work tool targets, and elevator operating component targets; The characteristic target trajectory includes the trajectory of the operator, the trajectory of the work tool, and the interaction trajectory of the operating components; In S02, the method for identifying worker targets, work tool targets, and elevator operating component targets based on the target detection model includes: The YOLOv8 target detection model is used to detect workers, wrenches, measuring tools, protective equipment and elevator components in video frames. The cross-frame target correspondence is established through a multi-target association tracking algorithm to generate target trajectory sequences.
4. The elevator inspection and maintenance operation supervision method based on spatiotemporal behavior recognition as described in claim 1, characterized in that, In S02, the method for extracting human key point information of workers based on the posture estimation model includes: The OpenPose posture estimation model is used to extract key points of the head, shoulders, elbows, wrists, hips, knees and / or ankles of the worker. Combined with the relative position changes of key points, joint angle changes and key point velocity changes, posture state features used to characterize the work posture are obtained.
5. The elevator inspection and maintenance operation supervision method based on spatiotemporal behavior recognition as described in claim 1, characterized in that, In S04, the key action semantic descriptor includes at least action category information, tool category information, active component information, contact state information, human posture state information and / or action duration information; The contact state information is determined by the spatial proximity and overlap relationships between the tool target area, the key point area of the human hand, and the operating component area.
6. The elevator inspection and maintenance operation supervision method based on spatiotemporal behavior recognition as described in claim 1, characterized in that, In S05, the semantic graph of work behavior at least represents the relationship between workers, work tools, operating components, scene areas, action events and equipment status; In S05, the nodes in the semantic graph of the work behavior include personnel nodes, tool nodes, component nodes, area nodes, action nodes and / or status nodes; The edges in the semantic graph of the job behavior include edges related to temporal sequence, contact, spatial adjacency, action dependency, and / or state constraint.
7. The elevator inspection and maintenance operation supervision method based on spatiotemporal behavior recognition as described in claim 1, characterized in that, In S06, the standard operation template is pre-set with one or more of the following: a set of prescribed actions for the corresponding operation item, a prescribed sequence of actions, a correspondence between tools and parts, a range of allowed duration of actions, human safety posture constraints, and elevator operation status constraints. The task judgment results include one or more of the following: task continuity result, action sequence result, tool usage consistency result, posture standardization result, and action completion result; The compliance determination result includes at least one of the following: normal, missing items, out of order, misuse of tools, abnormal posture, and incomplete key actions.
8. The elevator inspection and maintenance operation supervision method based on spatiotemporal behavior recognition as described in any one of claims 1 to 7, characterized in that, It also includes: S07. Based on the compliance assessment results and risk level, output standardized behavior labels, risk warning information and / or traceability evidence data.
9. The elevator inspection and maintenance operation supervision method based on spatiotemporal behavior recognition as described in claim 8, characterized in that, The traceability evidence data includes keyframe images, key action timestamps, standardized behavior tag sequences, compliance scoring results, risk event records, and / or index information corresponding to the original video clips.
10. An elevator inspection and maintenance operation monitoring system based on spatiotemporal behavior recognition, characterized in that, It includes: The multi-source data acquisition module is used to collect multi-view video streams, work task data, elevator operating status data, and equipment component identification data from the elevator inspection and maintenance work area. The data preprocessing module is used to perform time synchronization, task binding, and candidate action segmentation on the multi-view video stream, task data, elevator operation status data, and equipment component identification data. The target parsing module is used to identify targets such as workers, tools, and elevator operating components, and to extract key information about the workers' bodies. The spatiotemporal behavior recognition module is used to extract static spatial features and dynamic temporal features from candidate action segments based on the I3D model and the SlowFast model, and obtain spatiotemporal behavior feature vectors. The key action extraction module is used to extract key action semantic descriptors of candidate action segments based on inter-frame difference results, action recognition confidence change results, semantic segmentation results, human key point information and target contact relationship. The semantic graph construction module is used to construct a job behavior semantic graph based on the spatiotemporal behavior feature vector, key action semantic descriptors, and temporal sequence relationships. The assessment and early warning module is used to match the semantic graph of the work behavior with the standard operation template and output the compliance judgment result and risk level. The data archiving and output module is used to output standardized behavioral labels, risk warning information, and traceability evidence data.