Elevator dangerous behavior real-time alarm method based on computer vision technology

Through the improved YOLOv8 model and Simcc+Resnet network, combined with the color segmentation of elevator pedals, the problem of detecting abnormal behaviors in the elevator is solved, and the safety monitoring and alarm of elevators is realized.

CN120298648APending Publication Date: 2025-07-11HANGZHOU DAZHU YUNZHI TECH CO LTD

Patent Information

Application Number
CN202311462288.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art cannot effectively detect and prevent abnormal behaviors of people in elevators, such as retrograde, falls, etc., resulting in safety hazards.

Method used

Using a computer vision-based method, the improved YOLOv8 model is used for object detection, and the skeleton point estimation is performed in combination with the Simcc+Resnet network. By analyzing the skeleton point position and the color segmentation of the elevator pedal, dangerous behaviors in the elevator are identified and judged in real time, and alarms are made through the visual perception platform.

Benefits of technology

Real-time and accurate identification and alarm of personnel behavior in the elevator, ensuring the safe operation of the elevator and the safety of passengers, and improving the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298648A_ABST
    Figure CN120298648A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent security, and provides a computer vision technology-based elevator dangerous behavior real-time alarm method, which comprises the following steps of: capturing the activity condition of an elevator in real time through a camera, processing an RGB (Red, Green, Blue) image transmitted by the camera through an algorithm platform, and determining the accurate position of a person in the image through a target detection technology; and human body skeleton point detection is further performed according to the target detection result, and is compared with a preset elevator handrail area, so that actions such as head stretching, hand stretching and falling down are effectively recognized. Meanwhile, according to the moving direction of the human body detection frame, the system can judge whether the passenger runs reversely or not. Besides accurate identification and judgment of potential dangerous behaviors such as retrograde running, hand stretching and head stretching of passengers on the elevator, the method can also identify large luggage and elevator pedal gaps, and comprehensive application of various algorithms ensures operation safety of the elevator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent security technology, and in particular to a real-time alarm method for dangerous elevator behaviors based on computer vision technology. Background Art

[0002] Automatic escalators play an important role in public places such as shopping malls, subways, and railway station entrances and exits. However, during peak hours, uncivilized behaviors such as crowding and pushing are often seen, causing emergencies such as people falling. In order to prevent major safety accidents caused by people falling, the escalator must be stopped immediately when a person falls.

[0003] With the development of intelligent video technology, some studies have been conducted on the application of intelligent video technology to detect whether pedestrians have abnormal behaviors such as falling. The Chinese invention patent with the publication number "CN101695983A" discloses an escalator energy saving and installation monitoring system based on omnidirectional computer vision. It proposes a method of using video technology to solve the energy saving and safety monitoring of elevators, but fails to explain how to detect abnormal behaviors of people; the Chinese invention patent with the publication number "CN110294398A" discloses an escalator remote monitoring control system and method, and also fails to explain how to detect abnormal behaviors of people on the escalator. Therefore, it cannot be avoided that people on the escalator will be trampled due to retrograde or falling, posing a hidden danger to the personal safety of passengers. Summary of the invention

[0004] In view of the above problems, the present invention proposes a real-time alarm method for dangerous elevator behaviors based on computer vision technology, which is mainly aimed at smart elevator scenarios and consists of a camera that transmits data from the outside world and a visual perception platform that integrates multiple algorithms.

[0005] The present invention provides a real-time alarm method for dangerous elevator behaviors based on computer vision technology, comprising the following steps:

[0006] S1. Perform target detection on each frame of image input to the platform through the improved YOLOv8 model; the detection objects are people and large luggage. If a detection frame representing large luggage appears in the input image, the front-end interface of the algorithm platform prompts an alarm. If a pedestrian image appears, the image p in the pedestrian detection frame is obtained. i Intercept and carry out the next step of algorithm detection;

[0007] S2, the pedestrian image p i To estimate skeleton points, the network structure used by the algorithm is Simcc+Resnet, and the skeleton points are set to [s0s1s2…s 17 ];

[0008] S3. Among the extracted skeleton points, s4, s3, s7, and s6 represent the positions of the left palm, left elbow, right palm, and right elbow respectively. Based on these four points, it is determined whether the pedestrian reaches out.

[0009] S4. Among the extracted skeleton points, s0 represents the position of the person's nose. Based on this point, it is determined whether the pedestrian stretches their head.

[0010] S5. Among the extracted skeleton points, s9, s 10 , s 12 , s 13 , s2, s5. These six points represent the positions of the person's left knee, left foot sole, right knee, right foot sole, left shoulder, and right shoulder respectively. Based on these six points, it is determined whether the pedestrian falls.

[0011] S6. The direction of movement of the pedestrian detection bounding box is used to determine whether the pedestrian is going in the opposite direction.

[0012] S7. Specific color blocks are set on the elevator pedal, and color segmentation technology is used to detect the change of the elevator gap in real time to determine whether there is a gap in the pedal.

[0013] Preferably, in the YOLOv8 model, the "C2f" structure is adopted in the detection model backbone network and Neck, which contains more skip connections and additional Split operations to replace the "C3" structure in the YOLOv8 model. The PaFPN structure is used to construct the feature pyramid of YOLO. In the detection head (Head) part, a decoupled head is adopted, and the Anchor-Free method is used. Two parallel branches are used to extract class and location features, and a 1×1 convolution is used for each layer to complete the classification and localization tasks, and the coordinate attention (CA) module is introduced. The implementation of CA attention is divided into two parallel stages. The input feature map is globally average pooled in the width and height directions respectively to obtain the feature maps in the width and height directions respectively. Assuming the shape of the input feature layer is [C, H, W], after average pooling in the width direction, the shape of the obtained feature layer is [C, H, 1].

[0014] For example, the output expression of the c-th channel with height h is:

[0015]

[0016] Similarly, the output expression of the c-th channel with width w is

[0017]

[0018] The overall formula is as follows

[0019] M c(F) = σ(MLP(AvgPool(H)) + MLP(AvgPool(W)))

[0020] Then, the two parallel stages are merged, the width and height are transposed to the same dimension, and then stacked to merge the width and height features. At this time, the feature layer we obtain is: [C, 1, H + W]. We use convolution (CNN) + normalization (BatchNorm) + activation function (Non-linear) to obtain features, and then separate them into two parallel stages again. Then, the width and height are separated into: [C, 1, H] and [C, 1, W]. After that, they are transposed to obtain two feature layers [C, H, 1] and [C, 1, W]. Then, we use 1x1 convolution to adjust the number of channels and take sigmoid to obtain the attention situation in the width and height dimensions:

[0021] g h = σ(F h (f h ))

[0022] g w = σ(F w (f w ))

[0023] Multiplying by the original features is the feature map output after the CA attention mechanism, which is expressed as the following formula:

[0024]

[0025] Preferably, the specific calculation method for determining whether a pedestrian reaches out in step S3 is as follows:

[0026] On the perception platform, two lines l l , l r representing the escalator guardrail are pre-planned as marking lines, and l l , l r The input image is divided into three regions a l , a m , a r where a m represents the area where the elevator is located, and a l and a r represent the areas on both sides of the elevator. For the two points s4 and s3 of the left hand, calculate their distances d l and d ll to both sides of a lr If

[0027]

[0028] it means that the point is within the area of a l and the pedestrian does not stretch. Otherwise, it means that the point is outside the area and the pedestrian reaches out. The same applies to the two points of the right hand. If

[0029]

[0030] It is explained that the point is in area a l If the pedestrian does not stretch out their hand within the area, otherwise it means the point is outside the area, the pedestrian stretches out their hand, and the perception platform alarms.

[0031] Preferably, the specific calculation method for determining whether a pedestrian stretches their head in step S4 is as follows:

[0032] On the perception platform, two lines l l , l r represented by the escalator guardrail are pre-planned as marking lines, l l , l r The input picture is divided into three areas a l , a m , a r where a m represents the area where the elevator is located, and a l and a r represent the areas on both sides of the elevator. For the nose point s0, calculate the distances d m and d ml from it to both sides of a mr If

[0033]

[0034] It is explained that the point is in area a m If the pedestrian does not stretch their head within the area, otherwise it means the point is outside the area, the pedestrian stretches their head, and the perception platform alarms.

[0035] Preferably, the specific method for determining whether a pedestrian falls in step S5 is: Through the input skeleton points s9, s 10 , s 12 , s 13 , s2, s5, calculate the angles between the straight lines connecting the left sole and the left shoulder with the horizontal line, the left sole and the left knee with the horizontal line, the right sole and the right shoulder with the horizontal line, and the right sole and the right knee with the horizontal line in real time. Set a threshold of about plus or minus 10°. If the difference between two or more of the four angles and 90° exceeds this threshold, it is determined that the pedestrian has fallen, and the perception platform alarms.

[0036] Preferably, when multiple pedestrians appear in the pedestrian detection frame, the following algorithm is used to achieve real-time target tracking of pedestrians: When a pedestrian is first detected, their position is directly given by the detection frame, the speed is initialized to 0, and the state prediction formula is:

[0037] x pred = x + dx × Δt

[0038] y pred =y+dy×Δt

[0039] Where Δt is the time interval between two frames. When the detection box corresponding to the new frame is found, the state is updated as follows:

[0040]

[0041]

[0042] where x current and current is the center of a pedestrian box in the current frame, x previous and previous is the predicted position of the previous frame. In order to associate targets in consecutive frames, assuming that there are M targets in the previous frame and N targets in the current frame, in order to find the best match, we first need to calculate an M×N cost matrix, where each element cost(i,j) represents the cost between the i-th target in the previous frame and the j-th target in the current frame.

[0043] Typically, this cost is the Euclidean distance between two targets, if x i and i is the coordinate of the i-th target in the previous frame, and x j and j is the coordinate of the jth target in the current frame, then:

[0044]

[0045] The Hungarian algorithm is used to find the best target match of the previous frame for the target of the current frame to minimize the overall cost. The above method can be used to achieve real-time and accurate target tracking of pedestrians on the elevator. After tracking, the real-time moving trajectory of the center point of the same pedestrian detection frame is compared with the direction of elevator operation. If it is opposite to the direction of elevator operation, it is judged to be reversed and the perception platform alarm is triggered.

[0046] Preferably, the step S7 is specifically as follows: First, stick the predetermined yellow color patches on the elevator pedal. In the H (hue) S (saturation) V (value) color space, the hue disk of yellow is between 11 - 34°. First, all the yellow colors in the input picture are segmented through the HSV color space. Then, set the area threshold, and only the color patches exceeding this threshold will be displayed, and the result is converted into a grayscale image. In this way, only the black background and white color patches will be left on the screen. The findContours function of OpenCV can be used to find the borders of the color patches, and then the centroid of the color patches can be found through the moments function. Connecting the two centroids can obtain the slope of this straight line. If there are gaps or displacements in the pedal, the slope of the straight line will change significantly. At this time, the perception platform will give an alarm.

[0047] Compared with the prior art, the present invention has the following advantages:

[0048] 1. The present invention introduces an improved YOLOv8 target detection model. This model not only significantly improves the recognition accuracy but also further optimizes the overall performance of the system on the basis of ensuring real-time performance, thereby enhancing the stability and reliability of the system.

[0049] 2. The present invention cleverly combines target detection to achieve accurate identification of large luggage and pedestrians, ensuring that the targets in the monitoring area are accurately captured and analyzed.

[0050] 3. Through in-depth research and analysis of the pedestrian skeleton points, the present invention successfully conducts detailed recognition and analysis of various behaviors of pedestrians. Using the angles formed by the skeleton points and their intersections with the predetermined area, it is possible to accurately determine high-risk behaviors in the elevator such as falling, reaching out, and stretching, and ensure the real-time performance and accuracy of this recognition process.

[0051] 4. The present invention provides a real-time monitoring and alarm mechanism to ensure the normal operation of the elevator and the safety of passengers. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is the flowchart of the elevator warning module of the visual perception platform of the present invention;

[0053] Figure 2 is the main interface of the visual perception platform of the present invention;

[0054] Figure 3 is the alarm interface of the visual perception platform of the present invention;

[0055] Figure 4 is the skeleton point diagram generated by the present invention;

[0056] Figure 5 is the improved YOLOv8 model diagram of the present invention;

[0057] Figure 6 This is the detailed structural diagram of each module of the improved YOLOv8 model network of the present invention. Specific implementation manners

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.

[0059] The present invention proposes a real-time alarm method for elevator dangerous behaviors based on computer vision technology, mainly for intelligent elevator scenarios. The present invention consists of a camera for transmitting external data and a visual perception platform integrated with various algorithms, and the specific functions are as follows.

[0060] First, the platform connects to the camera in front of the elevator through the RTSP protocol and completes the following algorithm detection through the RGB images input by the camera. Target detection is performed on each frame input to the platform through the improved YOLOv8 model, and the detection objects are people and large luggage. If a detection frame representing large luggage appears in the input picture, the front-end interface of the algorithm platform will prompt an alarm, and then the image p in the pedestrian detection frame obtained will be i cropped out for the next algorithm detection. The platform will obtain the pedestrian image p i for skeleton point estimation. The network structure used by the algorithm is Simcc + Resnet, and the obtained skeleton points are set as [s0 s1 s2…s 17 . Then, the obtained skeleton points are used for subsequent judgment of human behaviors such as falling, reaching out, and stretching the head. At the same time, the obtained human detection frame is used for the analysis of going against the current. At the same time, among the obtained skeleton points, s4, s3, s7, and s6 represent the positions of the left palm, left elbow, right palm, and right elbow respectively, and these 4 points are used as the basis for judging whether to reach out. Further, among the obtained skeleton points, s0 represents the position of the person's nose, which is used as the basis for judging whether to stretch the head. Among the obtained skeleton points, s9, s 10 , s 12 , s 13 , s2, and s5 these six points represent the positions of the left knee, left foot sole, right knee, right foot sole, left shoulder, and right shoulder of the person respectively, and these six points are used as the basis for judging whether to fall. Then, the direction of movement of the pedestrian detection frame is used to judge whether to go against the current. Finally, through the specific color blocks pre-set on the elevator pedal, the system uses color segmentation technology to detect the change of the elevator gap in real time to ensure the safety of the elevator.

[0061] Improved YOLOv8 object detection model, which adopts the "C2f" structure in the backbone network and Neck of the detection model, including more skip connections and additional Split operations, to replace the "C3" structure in the YOLOv8 model.

[0062] At the same time, the PaFPN structure is used to construct the feature pyramid of YOLO to fully fuse multi-scale information. In the detection head (Head) part, a decoupled head is adopted, and the Anchor-Free method is used. Two parallel branches are used to extract class and location features, and a 1×1 convolution is used for each layer to complete the classification and localization tasks. The present invention introduces a coordinate attention (CA) module. The implementation of CA attention can be considered to be divided into two parallel stages. The input feature map is globally average pooled in the width and height directions respectively to obtain feature maps in the width and height directions. Assuming that the shape of the input feature layer is [C, H, W], after average pooling in the width direction, the shape of the obtained feature layer is [C, H, 1].

[0063] For example, the output expression of the c-th channel with height h is:

[0064]

[0065] Similarly, the output expression of the c-th channel with width w is

[0066]

[0067] The overall formula is as follows

[0068] M c (F) = σ(MLP(AvgPool(H)) + MLP(AvgPool(W)))

[0069] Then, the two parallel stages are merged, the width and height are transposed to the same dimension, and then stacked to merge the width and height features. At this time, the obtained feature layer is: [C, 1, H + W]. Features are obtained using convolution (CNN) + batch normalization (BatchNorm) + activation function (Non-linear). Then, it is separated into two parallel stages again, and the width and height are separated into: [C, 1, H] and [C, 1, W], and then transposed. Two feature layers [C, H, 1] and [C, 1, W] are obtained. Then, the number of channels is adjusted using a 1x1 convolution and sigmoid is taken to obtain the attention situation in the width and height dimensions:

[0070] g h = σ(F h (f h ))

[0071] gw = σ(F w (f w ))

[0072] Multiplying by the original feature is the feature map output after the CA attention mechanism, which is expressed as the following formula:

[0073]

[0074] The present invention is trained on the collected dataset. The data collection method is to transmit the data back in real time by the camera, and one frame is taken every fixed number of frames. The Labelme software is used to label the pedestrians and large pieces of luggage with rectangular frames. The data is divided into a training set and a validation set according to a certain ratio, trained on the training set, and validated on the validation set. The model with the best performance on the validation set after the iteration ends is used for actual application.

[0075] The metric for the validation set is mAP0.5:0.95, and the calculation method of mAP is as follows::

[0076]

[0077] When the IoU between the detection box and the actual annotation is greater than a certain threshold, the detection result is considered correct. mAP(0.50:0.95) represents the average mAP value in different IoU threshold ranges (from 0.5 to 0.95, with a step size of 0.05).

[0078] AP is the area under the precision-recall curve. The calculation methods of precision and recall are as follows:

[0079]

[0080]

[0081] TN represents the true negative prediction, FP represents the false positive prediction, and FN represents the false negative prediction.

[0082] Among the extracted skeleton points, s4, s3, s7, and s6 represent the positions of the left hand palm, left hand elbow, right hand palm, and right hand elbow respectively. These 4 points are used as the basis for judging whether to reach out. The calculation method is as follows:

[0083] On the perception platform, two lines l l , l r are pre-planned as the marking lines, l l , lr Divide the input image into three regions a l , a m , a r where a m represents the region where the elevator is located, a l and a r represent the regions on both sides of the elevator. For the two points s4 and s3 on the left hand, calculate their distances d l and d ll to both sides of a lr If

[0084]

[0085] it indicates that the point is within the region a l and the pedestrian has no stretching. Otherwise, it means the point is outside the region and the pedestrian reaches out. The same applies to the two points on the right hand. If

[0086]

[0087] it indicates that the point is within the region a l and the pedestrian has no reaching out. Otherwise, it means the point is outside the region and the pedestrian reaches out, and the sensing platform alarms.

[0088] Among the extracted skeleton points, s0 represents the position of the person's nose. This point is used as the basis for judging whether the person stretches their head. The calculation method is as follows:

[0089] Pre-plan two lines l l , l r on the sensing platform to represent the escalator guardrail as marking lines, l l , l r Divide the input image into three regions a l , a m , a r where a m represents the region where the elevator is located, a l and a r represent the regions on both sides of the elevator. For the nose point s0, calculate its distances d m and d ml to both sides of a mr If

[0090]

[0091] it indicates that the point is within the region a m and the pedestrian has no head stretching. Otherwise, it means the point is outside the region and the pedestrian stretches their head, and the sensing platform alarms.

[0092] Among the extracted skeleton points, s9, s 10 , s 12 , s 13, The six points s2 and s5 respectively represent the positions of a person's left knee, left foot sole, right knee, right foot sole, left shoulder, and right shoulder. When a person stands normally, the straight line connecting the left foot sole and the left shoulder should form a right angle with the horizontal line, and the straight line connecting the left foot sole and the left knee should also be a right angle. The same applies to the right side. Utilizing this phenomenon, the present invention calculates in real time the angles between the straight lines connecting the left foot sole and the left shoulder, the left foot sole and the left knee, the right foot sole and the right shoulder, and the right foot sole and the right knee with the horizontal line based on the input skeleton points, sets a threshold of about plus or minus 10°, and if the difference between two or more of the four angles and 90° exceeds this threshold, it is determined that the pedestrian has fallen, and the perception platform issues an alarm.

[0093] Multiple pedestrians will appear simultaneously in the input images. Therefore, to achieve real-time one-to-one correspondence between pedestrians and detection frames. The present invention realizes real-time object tracking of pedestrians through the following algorithm:

[0094] When a pedestrian is first detected, its position is directly given by the detection frame, and the speed is initialized to 0. The state prediction formula is:

[0095] x pred = x + dx × Δt

[0096] y pred = y + dy × Δt

[0097] Where Δt is the time interval between two frames. When the detection frame corresponding to the new frame is found, the state is updated as follows:

[0098]

[0099]

[0100] Where x current and y current are the centers of a certain pedestrian frame in the current frame, and x previous and y previous are the predicted positions in the previous frame. To associate targets in consecutive frames, assume we have M targets in the previous frame and N targets in the current frame. To find the best match, we first need to calculate an M×N cost matrix, where each element cost(i,j) represents the cost between the i-th target in the previous frame and the j-th target in the current frame.

[0101] Usually, this cost is the Euclidean distance between two targets. If x i and y i are the coordinates of the i-th target in the previous frame, and x j and y jis the coordinate of the jth target in the current frame, then:

[0102]

[0103] Use the Hungarian algorithm to find the best previous frame target match for the current frame target to minimize the overall cost.

[0104] The above method can be used to achieve real-time and accurate target tracking of pedestrians on the elevator. After tracking is achieved, the real-time moving trajectory of the center point of the same pedestrian detection frame is compared with the direction of elevator operation. If it is opposite to the direction of elevator operation, it is judged as reverse travel and the sensing platform alarm is triggered.

[0105] Finally, a predetermined yellow color block is pasted on the elevator pedal. The yellow hue disk is between 11-34° in the H (hue) S (saturation) V (brightness) color space. The present invention first separates all the yellow in the input image through the HSV color space. Then set the area threshold, only the color blocks exceeding the threshold will be displayed, and the result will be converted to a grayscale image, so that only the black background and white color blocks will be left on the screen. The border of the color block can be found using OpenCV's findContours function, and then the center of mass of the color block can be found using the moments function. The slope of the straight line can be obtained by connecting the two center of mass. If there is a gap or movement in the pedal, the slope of the straight line will change significantly, and the perception platform will alarm at this time.

[0106] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. It should therefore be understood that many modifications may be made to the exemplary embodiments and that other arrangements may be devised without departing from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in a manner different from that described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be used in other described embodiments.

Claims

1. A real-time alarm method for elevator dangerous behaviors based on computer vision technology, characterized in that, Including the following steps: S1. Perform object detection on each frame of the image input to the platform using an improved YOLOv8 model; the detection objects are people and large luggage. If a detection box representing large luggage appears in the input image, the front-end interface of the algorithm platform will prompt an alarm. If a pedestrian image appears, the image p in the obtained pedestrian detection box will be i cropped out for the next algorithm detection; S2. Estimate the skeleton points for the obtained pedestrian image p i The network structure used by the algorithm for skeleton point estimation is Simcc+Resnet, and the obtained skeleton points are set as [s0 s1 s2…s 17 ; S3. Among the extracted skeleton points, s4, s3, s7, and s6 represent the positions of the left palm, left elbow, right palm, and right elbow respectively. Based on these four points, it is determined whether the pedestrian reaches out his hand. S4. Among the extracted skeleton points, s0 represents the position of the person's nose. Based on this point, it is determined whether the pedestrian stretches his head. S5. Among the extracted skeleton points, s9, s 10 , s 12 , s 13 , s2, s5. These six points respectively represent the positions of a person's left knee, left foot sole, right knee, right foot sole, left shoulder and right shoulder. Based on these 6 points, it is determined whether the pedestrian has fallen; S6. The direction of movement of the pedestrian detection frame is used to determine whether the pedestrian is going against the flow. S7. Specific color blocks are set on the elevator pedal, and the change of the elevator gap is detected in real time through color segmentation technology to determine whether there is a gap in the pedal.

2. The real-time alarm method for elevator dangerous behaviors based on computer vision technology according to claim 1, characterized in that In the YOLOv8 model, the "C2f" structure is adopted in the detection model backbone network and Neck, which contains more skip connections and additional Split operations to replace the "C3" structure in the YOLOv8 model. The PaFPN structure is used to construct the feature pyramid of YOLO. In the detection head (Head) part, a decoupled head is adopted, and the Anchor-Free method is used. Two parallel branches are used to extract category and location features, and a 1×1 convolution is used for each layer to complete the classification and positioning tasks. The coordinate attention (CA) module is introduced. The implementation of CA attention is divided into two parallel stages. The input feature map is globally average pooled in the width and height directions respectively to obtain the feature maps in the width and height directions respectively. Assuming the shape of the input feature layer is [C, H, W], after the average pooling in the width direction, the shape of the obtained feature layer is [C, H, 1]. For example, the output expression formula of the c-th channel with height h is: Similarly, the output expression formula of the c-th channel with width w is The overall formula is as follows M c (F) = σ(MLP(AvgPool(H)) + MLP(AvgPool(W))) Then the two parallel stages are merged, the width and height are transposed to the same dimension, and then stacked to merge the width and height features. At this time, the feature layer we obtain is: [C, 1, H+W]. Features are obtained by using convolution (CNN)+batch normalization (BatchNorm)+activation function (Non-linear). Then it is separated into two parallel stages again, and the width and height are separated into: [C, 1, H] and [C, 1, W]. Then transpose to obtain two feature layers [C, H, 1] and [C, 1, W]. Then use a 1x1 convolution to adjust the number of channels and take the sigmoid to obtain the attention situation in the width and height dimensions: g h = σ(F h (f h )) g w = σ(F w (f w )) Multiplying by the original feature is the feature map output after the CA attention mechanism, as expressed in the following formula:

3. A real-time alarm method for elevator dangerous behaviors based on computer vision technology according to claim 1, characterized in that, The specific calculation method for determining whether a pedestrian reaches out his hand in step S3 is as follows: Plan in advance the two lines l represented by the escalator guardrail on the perception platform l , l r As marking lines, l l , l r Divide the input image into three regions a l , a m , a r Among them, a m Represents the area where the elevator is located, a l And a r Represent the areas on both sides of the elevator. For the two points s4 and s3 on the left hand, calculate their distances d l To both sides of a ll And d 1r If The description point is at a l Within the area, the pedestrian does not stretch. Conversely, it means the point is outside the area and the pedestrian reaches out. The same applies to the two points of the right hand. If The description point is at a l Within the area, the pedestrian does not extend their hand. Conversely, if the point is outside the area and the pedestrian extends their hand, the perception platform will alarm.

4. A real-time alarm method for elevator dangerous behaviors based on computer vision technology according to claim 1, characterized in that The specific calculation method for determining whether a pedestrian stretches his head in step S4 is as follows: Plan in advance the two lines \(l\) represented by the escalator guardrail on the perception platform l ,\(l\) r As marking lines,\(l\) l ,\(l\) r Divide the input image into three regions \(a\) l ,\(a\) m ,\(a\) r Among them,\(a\) m represents the region where the elevator is located,\(a\) l and \(a\) r represent the regions on both sides of the elevator. Calculate the distances \(d\) m and \(d\) ml from the nose point \(s0\) to both sides of \(a\) mr If It shows that the point is at a m Within the area, the pedestrian does not stretch their head. Conversely, it indicates that the point is outside the area, the pedestrian stretches their head, and the perception platform gives an alarm.

5. A real-time alarm method for elevator dangerous behaviors based on computer vision technology according to claim 4, characterized in that The specific method for determining whether a pedestrian has fallen in step S5 is as follows: Based on the input skeleton points s9, s 10 , s 12 , s 13 , s2, and s5, calculate in real time the angles between the lines connecting the left sole and the left shoulder, the left sole and the left knee, the right sole and the right shoulder, and the right sole and the right knee with the horizontal line. Set a threshold of about plus or minus 10°. If the difference between two or more of the four angles and 90° exceeds this threshold, it is determined that the pedestrian has fallen, and the perception platform gives an alarm.

6. The real-time alarm method for elevator dangerous behaviors based on computer vision technology according to claim 4, wherein, If there are multiple pedestrians in the pedestrian detection frame, the following algorithm is used to achieve real-time object tracking of pedestrians: When a pedestrian is first detected, its position is directly given by the detection frame, and the speed is initialized to 0. The state prediction formula is: x pred = x + dx × Δt y pred = y + dy × Δt where Δt is the time interval between two frames. When the detection frame corresponding to the new frame is found, the state is updated as follows: where x current and y current are the centers of a pedestrian bounding box in the current frame, x previous and y previous are the predicted positions in the previous frame. To associate targets across consecutive frames, assume there are M targets in the previous frame and N targets in the current frame. To find the best match, first, a cost matrix of size M×N needs to be calculated, where each element cost(i, j) represents the cost between the i-th target in the previous frame and the j-th target in the current frame. Typically, this cost is the Euclidean distance between two targets. If x i and y i are the coordinates of the i-th target in the previous frame, and x j and y j are the coordinates of the j-th target in the current frame, then: The Hungarian algorithm is used to find the best target match of the previous frame for the target of the current frame to minimize the overall cost. The above method can be used to achieve real-time and accurate target tracking of pedestrians on the elevator. After tracking, the real-time moving trajectory of the center point of the same pedestrian detection frame is compared with the direction of elevator operation. If it is opposite to the direction of elevator operation, it is judged to be reversed and the perception platform alarm is triggered.

7. A real-time alarm method for elevator dangerous behaviors based on computer vision technology according to claim 1, characterized in that, The step S7 is specifically as follows: firstly, a predetermined yellow color block is pasted on the elevator pedal. In the H (hue) S (saturation) V (brightness) color space, the yellow hue wheel is between 11-34°. Firstly, all the yellow in the input image is segmented out through the HSV color space. Then, an area threshold is set. Only the color blocks exceeding the threshold are displayed. The result is converted into a grayscale image. In this way, only a black background and a white color block are left on the screen. The border of the color block can be found by using the findContours function of OpenCV. Then, the center of mass of the color block can be found by using the moments function. The slope of the straight line can be obtained by connecting the two center of mass. If there is a gap or movement in the pedal, the slope of the straight line will change significantly. At this time, the sensing platform will alarm.

Citation Information

Patent Citations

  • Omnibearing computer vision based energy-saving and safety monitoring system of escalator

    CN101695983A

  • Escalator remote monitoring control system and method

    CN110294398A

Cited By

  • Depth learning-based armrest-unholding behavior detection system

    CN121811488A