Construction elevator standard operation monitoring and alarm system and method based on video analysis

Through the construction elevator monitoring system based on video analysis, the video analysis platform and cross-level integrated SSD detection framework are used to analyze the video stream of the construction elevator, and automatically identify people and objects, solving the problem of the existing technology being difficult to monitor other illegal operating status except overloading, and achieving efficient and accurate monitoring and alarms.

CN114529868BActive Publication Date: 2025-05-23CHINA CONSTR FIRST DIV GROUP CONSTR & DEV +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210071680.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-21
Publication Date
2025-05-23
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

The existing construction elevator monitoring and alarm technology is difficult to effectively monitor other illegal operating conditions except overload, such as carrying illegal items, and overload alarms are difficult to directly trace the direct cause of overload.

Method used

The construction elevator standard operation monitoring and alarm system is adopted based on video analysis, and video streams are collected through a webcam, and the video analysis platform and cross-level integrated SSD detection framework are used to analyze the video streams in real time, automatically identify people and objects, determine whether they comply with the construction elevator safety operation specifications, and generate relevant alarm signals.

Benefits of technology

Automatic identification of the types and quantities of construction elevator loads is realized, and other illegal operating status can be effectively monitored except overload, improving the accuracy and automation of monitoring, and reducing loopholes in manual supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529868B_ABST
    Figure CN114529868B_ABST
Patent Text Reader

Abstract

The present invention proposes a construction elevator standard operation monitoring and alarm system and method based on video analysis. The network camera in the construction elevator car is a collection device for the elevator monitoring video; the video stream transmission module is mainly composed of a WIFI communication module, which transmits the data collected by the network camera to the server; the server is located at the project site or in the cloud, and is mainly used to run the video analysis platform and provide an alarm interface; the video analysis platform has the main function of performing real-time analysis on the collected video stream and automatically identifying people and typical objects in the shooting field of view. The present invention uses computer vision technology to analyze the construction elevator monitoring video, automatically identify the load category and quantity, determine whether the current operating status meets the construction elevator safety operation specifications, and generate relevant alarm signals. The system of the invention is easy to deploy without the support of a large number of distributed hardware devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image analysis, and in particular relates to a construction elevator standard operation monitoring and alarm system based on video analysis and a method thereof. Background Art

[0002] In recent years, with the continuous expansion of my country's construction industry, safety accidents have occurred frequently, and the safety of construction sites has received great attention. Since the current construction industry involves a large number of high-altitude operations, such as the construction of high-rise buildings, the construction of three-dimensional intersection projects such as bridges, etc., construction elevators used to carry people and goods on construction sites are widely used. Compared with ordinary elevators, the operation purpose and operating environment of construction elevators are not ideal. Once abnormal behaviors such as falling occur, it is very easy to cause major safety accidents. In order to ensure the normal order of construction sites and avoid injuries, construction elevators have strict control over the number and type of people and goods in the car during operation. According to the abnormal number or type of loads, the monitoring alarm of the standardized operation status of construction elevators can be divided into two categories: overload alarm and violation alarm.

[0003] At present, the common practice of overload alarm is to use a weight sensor to monitor the real-time load of the elevator and send an alarm signal when it is overloaded. This method can be implemented simply and can achieve good results. However, it cannot monitor the objective factors that constitute overload, such as personnel overload or item overload. To solve the problem of personnel overload, the number of personnel in the elevator load can be counted by equipping personnel with radio frequency cards, as described in Chinese patent CN103708315. However, it requires the support of relevant hardware, and it is difficult to spread all possible loads on the construction site, and it is difficult to deal with emergencies such as illegal passengers who do not wear radio frequency modules. As for illegal alarms, it currently mainly relies on the supervision and control of on-site personnel, including patrolling the loading site of the construction elevator and manually checking the video taken by the surveillance camera. The reliability of the manual method depends on the execution strength of the supervisor, and there are problems such as large workload and easy loopholes. There are also methods that propose intelligent early warning of construction elevators through image recognition technology, as described in Chinese patent CN202414916U. However, this utility model patent only proposes the overall structure of an intelligent early warning system for a construction elevator, and provides a functional description of each module required to implement the early warning, but does not provide a detailed description of the specific image analysis method used by the video analysis module.

[0004] In summary, the current construction elevator monitoring and alarm technology is difficult to warn of other illegal operating conditions besides overloading, such as carrying illegal items. At the same time, the current technology for handling overloading alarms is relatively rudimentary, and it is difficult to directly trace the direct cause of overloading. By searching the currently authorized or published patents, no automated solution that can solve the above problems at the same time has been found. Summary of the invention

[0005] In view of the deficiencies of the prior art, the present invention proposes a construction elevator specification operation monitoring and warning system and method based on video analysis. The purpose of the present invention is to overcome the defects of the existing construction elevator warning technology, use computer vision technology to analyze the monitoring video of the construction elevator, automatically identify the load category and quantity, judge whether the current operation state conforms to the safety operation specifications of the construction elevator, and generate relevant warning signals.

[0006] The specific technical solution is as follows:

[0007] A construction elevator specification operation monitoring and warning system based on video analysis includes a network camera installed in the construction elevator car, a video stream transmission module and a server, a video analysis platform and a warning interface;

[0008] The network camera in the construction elevator car is a collection device for the elevator monitoring video;

[0009] The video stream transmission module, mainly composed of a WIFI communication module, transmits the data collected by the network camera to the server;

[0010] The server, located at the project site or in the cloud, is mainly used to run the video analysis platform and provide a warning interface;

[0011] The video analysis platform, whose main function is to perform real-time analysis on the collected video stream and automatically identify the personnel and typical objects in the shooting field of view.

[0012] Using the above-mentioned construction elevator specification operation monitoring and warning system based on video analysis, it includes the following steps:

[0013] S1. The video analysis platform frames the input video stream, extracts image frames by equidistant sampling, and preprocesses the above image frames to align the color space to obtain preprocessed image frames;

[0014] S2. Use a cross-level fusion SSD detection framework to perform target detection on the preprocessed image frames obtained by frame sampling. The target detection environment is a construction elevator, and the detection targets include items not allowed to enter the construction elevator, elevator passengers, etc.;

[0015] S3. Judge whether there is an abnormal behavior that violates the safety specifications according to the judgment conditions;

[0016] The judgment conditions are: the objects detected based on the target in step S2 do not belong to the items allowed to enter the construction elevator; the number of detected passengers exceeds the allowed number; if the above conditions are met, it is determined that a violation has occurred, the warning interface is called, and a warning signal is output, otherwise return to step S1 for the next detection.

[0017] Preferably, S1 is specifically performed as follows: using some preprocessed image frames as training set image frames, first converting the training set image frames from RGB image space to Lab space, and calculating the mean of the training set image frames in the three channels of Lab space respectively. and variance m L , m a and m b Refers to the mean pixel value of L channel, a channel and b channel respectively, var L , var a and var b They refer to the pixel value variance of L channel, a channel and b channel respectively;

[0018] Then, the training set image frames are calculated pixel by pixel according to the three channels of L, a and b i∈{L,a,b}, where p i,x ,,m i,x and var i,x Respectively represent the pixel value, mean and variance of the pixel x in the current image in channel i; is the adjusted pixel value of pixel point x in channel i;

[0019] Finally, the image composed of the adjusted pixel values ​​is inversely transformed from the Lab space back to the RGB space to obtain the preprocessed image frame.

[0020] Preferably, the detection method of the cross-level fusion SSD detection framework in S2 comprises the following specific steps:

[0021] (1) Extracting target features

[0022] Aiming at the objective in claim 2, a cross-level fusion SSD algorithm is used to extract features from the training set image frames. The cross-level fusion SSD algorithm not only includes the extraction of deep feature maps by the traditional SSD algorithm, but also adds the extraction of two shallow feature maps on the basis of the traditional SSD algorithm. The sizes of the two newly added shallow feature maps are 150×150 and 76×76.

[0023] (2) Feature Fusion

[0024] First, multiple groups of cross-level feature maps are fused based on the two-stream convolutional model. The two cross-level feature maps to be fused in each group are recorded as shallow feature maps f shallow And the deep feature map f deep , the shallow feature map f shallow are the two shallow feature maps in step (1), and the deep feature map f deep is the deep feature map in step (1); where f shallow The size of f is larger, carrying more local features, deepThe size of is smaller and carries more global features; the fusion steps are specifically as follows: first, the deep feature map f is deconvolved by a deconvolution operation deep The size of the shallow feature map f shallow Consistent; then calculate the deep feature map f deep Channel outer product B(f shallow ,TransConv(f deep ))=f shallow T TransConv(f deep ) obtains the result f of the bilinear mapping; compresses the dimension of the fusion feature of f through multi-level dense convolution and pooling operations based on difference supervision to obtain a fusion feature map of appropriate size; the method uses a three-stage series convolution pooling module to compress the dimension of the result of the channel outer product; uses a difference supervision method to deconvolve the output of the high-level convolution pooling module, aligns the size to the adjacent low-level convolution pooling module, calculates the difference between the outputs of the two-level modules, and then performs a hole convolution on the difference to align the size to the fusion feature map to be output; performs inter-level difference supervision on the three-stage series convolution pooling modules respectively, and outputs the fusion feature;

[0025] (3) Dense prior box generation

[0026] The fused feature map obtained in step (2) is used to generate the prior frame, and the six sizes are: 38×38, 19×19, 10×10, 5×5, 3×3 and 1×1. Each n×n feature map has n×n center points, and each center point generates k prior frames. The k generated by each center point of each feature map in each layer of the six layers represented by the above six sizes are 4, 6, 6, 6, 4, and 4 respectively;

[0027] (4) Obtain the type and number of targets in the image

[0028] Classify and detect the a priori frame obtained in step (3) to obtain the type and number of targets in the current input image; each a priori frame will be used to perform the following tasks: classify the a priori frame through the softmax function to determine whether the a priori frame contains the target in claim 2, and if so, first determine the type of target, and then determine the number of targets of different types; adjust the size and position of the a priori frame through regression analysis to improve the accuracy and efficiency of target classification and detection;

[0029] (5) Training

[0030] For the proposed cross-level fusion SSD detection framework, input the training set image frames and loop through steps (1)-(4) to train the model until the target detection accuracy reaches the accuracy required by the actual project. The loop ends and the trained cross-level fusion SSD detection model is output.

[0031] Among them, in step (2), the specific feature map fusion is: fusing the 150×150 shallow feature map with the 19×19 deep feature map, and fusing the 76×76 shallow feature map with the 38×38 deep feature map.

[0032] Wherein, step (5) includes the following sub-steps:

[0033] (a) Construction of a monitoring video database of the construction elevator on site: First, a monitoring database based on the project is constructed, which mainly includes the on-site video data captured by the monitoring cameras installed in the construction elevator. The data size is more than 10G, and the number of various targets included should be as balanced as possible. The specific method is to continuously collect elevator monitoring videos for about one week and divide them into no less than 20,000 image frames; calibrate them manually, draw the position of the target in the image and give the category to which the target belongs;

[0034] (b) The proposed cross-level fusion SSD detection framework is trained. The objective function is obtained by weighting the prior box confidence loss and position loss. The specific form is: Use Adam optimizer for training; when the training process is fully converged, save the detection framework;

[0035] (c) When deploying at multiple project sites, repeat steps (a)-(b) according to the specific project, and train the detection framework parameters for the specific project site through fine-tuning.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] (1) Reduce the size of the feature map, thereby reducing the number of prior boxes. This operation helps improve the feature extraction network's ability to express small objects;

[0038] (2) Align the size of the fused feature map to the deep features fused with it through convolution and pooling operations Figure 1 Avoid introducing too much computational load;

[0039] (3) Implementing this operation in the image preprocessing stage solves the impact of large light fluctuations on image target recognition in industrial elevators and improves the accuracy of the target detection framework;

[0040] (4) In order to enable the detector to learn more actual scenes encountered by industrial elevators in engineering practice, the world's most widely used face detection benchmark dataset WIDER FICE is used; the engineering hat dataset contains about 10,000 photos of engineering hats in different actual situations; the engineering wooden stick dataset contains about 5,000 photos and the wheelbarrow dataset for pushing mud in engineering contains about 5,000 photos. As described above, the present invention can better meet the requirements of construction elevator target recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a system framework diagram of the present invention;

[0042] Figure 2 The video analysis platform workflow of the present invention;

[0043] Figure 3 This is a structural diagram of the cross-level fusion SSD detection framework used in the present invention;

[0044] Figure 4 The present invention provides a feature cross-level fusion method based on difference supervision;

[0045] Figure 5 Color normalization flow chart of the present invention. DETAILED DESCRIPTION

[0046] The specific technical solution of the present invention is explained in conjunction with embodiments.

[0047] like Figure 1 As shown, the construction elevator standard operation monitoring and alarm system based on video analysis includes a network camera arranged in the construction elevator car, a video stream transmission module and server, a video analysis platform and an alarm interface;

[0048] The network camera in the construction elevator car is a collection device for elevator monitoring video, and you can choose any brand on the market.

[0049] The video stream transmission module, mainly composed of a WIFI communication module, transmits the data collected by the network camera to the server;

[0050] Server, located at the project site or in the cloud, is mainly used to run the video analysis platform and provide an alarm interface;

[0051] The video analysis platform mainly performs real-time analysis on the collected video streams and automatically identifies people and typical objects in the shooting field of view.

[0052] The workflow of the construction elevator standard operation monitoring and alarm system based on video analysis is as follows Figure 2 As shown, the steps are:

[0053] S1, divide the input video stream into frames, extract image frames by equally spaced sampling, and perform preprocessing operation for color space alignment;

[0054] like Figure 5 As shown in Figure 1, the specific approach is: first, convert all images in the training set from RGB image space to Lab space, and calculate the mean of the three channels of the images in the training set in Lab space respectively. and variance

[0055] Then, all images are calculated pixel by pixel by channel where p i,x ,,m i,x and var i,x Represent the pixel value, mean and variance of the current image respectively is the pixel value after adjustment;

[0056] Finally, the image is transformed from Lab space back to RGB space; this operation can ensure the effectiveness of subsequent detection.

[0057] S2, adopt Figure 3 The cross-level fusion SSD detection framework shown performs target detection on the video frames obtained by frame sampling.

[0058] S3, determine whether there is any abnormal behavior that violates safety regulations;

[0059] The judgment conditions are: the detected object does not belong to the construction elevator access items; the detected passengers exceed the access number. If the above conditions are met, it is determined that a violation has occurred, the alarm interface is called, and an alarm signal is output. Otherwise, return to step S1 and perform the next detection.

[0060] The core algorithm of the video analysis platform adopted by the present invention is the cross-level fusion SSD detection framework. In order to improve the detection accuracy, the present invention improves the original SSD and better describes small targets by fusing shallow feature maps. Figure 3 As shown, the specific steps can be described as:

[0061] (1) Extract shallow features of the target

[0062] The traditional SSD uses the VGG model as the feature extraction backbone network. The specific feature graphs used are as follows: Figure 3 As shown by the light-colored arrows, the sizes are 38×38, 19×19, 10×10, 5×5, 3×3 and 1×1. In order to improve the ability to detect multi-scale targets, especially the ability to detect small targets, on this basis, the improvement method of the present invention is mainly to increase the shallow feature map, and improve the expression ability of the model by fusing the deep and shallow features. Figure 3As shown by the dark arrows in the middle, the sizes of the newly added feature maps are 150×150 and 76×76.

[0063] (2) Feature fusion. As mentioned above, the cross-level fusion SSD detection proposed in the present invention introduces shallow features with large size. If the prior box is directly generated, it will greatly increase the computational burden and affect the execution speed of the algorithm. Considering the balance between computational load and detection effect, this scheme first fuses multiple cross-level feature maps based on the two-stream convolution model. The fusion method is as follows: Figure 4 As shown. The two feature maps to be fused are respectively recorded as shallow features f shallow and deep features f deep , where f shallow The size of f is larger, carrying more local features, deep The size of is smaller and carries more global features. First, f deep The size of f is transformed into shallow Then calculate the channel outer product B(f shallow ,TransConv(f deep ))=f shallow T TransConv(f deep ), this operation can describe the features from the second-order statistics, thereby increasing the expressive power of the model. However, its essence is the outer product of the matrix, which will expand the feature dimension. In order to improve the computational efficiency of the model, the dimension of the fused feature is compressed through multi-level dense convolution and pooling operations based on difference supervision to obtain a fused feature map of appropriate size. Figure 4 As shown in the figure, a three-stage series of convolutional pooling modules is used to compress the dimension of the channel outer product result. In order to ensure the quality of the compressed features, a difference supervision method is used to deconvolve the output of the high-level convolutional pooling module, align the size to the adjacent low-level convolutional pooling module, calculate the difference between the outputs of the two-level modules, and then perform a dilated convolution on the difference to align the size to the fusion feature map to be output. The three-stage series of convolutional pooling modules are supervised by inter-level differences, and the output is the fusion feature.

[0064] In the present invention, the specific fusion level is: 150×150 shallow feature map is fused with 19×19 deep feature map, 76×76 shallow feature map is fused with 38×38 deep feature map, and the size of the fused feature map is aligned to the corresponding deep feature map, thereby reducing the size of the feature map and reducing the number of prior boxes. This operation helps to improve the expression ability of the feature extraction network for small targets.

[0065] (3) Dense prior box generation. Figure 3As shown in the figure, the feature maps used for prior frame generation in the improved framework are still of six sizes: 38×38, 19×19, 10×10, 5×5, 3×3 and 1×1. Due to the addition of the feature fusion module, the number of channels corresponding to 38×38 and 19×19 is different from that of the original framework. Each n×n feature map has n×n center points, and each center point generates k prior frames. The k generated by each center point in each of the six layers are 4, 6, 6, 6, 4, and 4 respectively.

[0066] (4) Classify and detect based on the prior frame to obtain the type and number of targets in the current input image. Each prior frame will be used to perform the following tasks: classify the prior frame through the softmax function to determine whether it contains the target; adjust the size and position of the prior frame through regression analysis.

[0067] (5) Training. The steps can be described as follows: (a) Construction of a surveillance video database of elevators on the project site. Considering that the elevator cars, lighting conditions, environmental factors, etc. of different project sites are different, we first build a project-based monitoring database, which mainly includes on-site video data taken by surveillance cameras installed in construction elevators. The data size is more than 10G, and the number of various targets included should be as balanced as possible. The specific method is to continuously collect elevator surveillance videos for about one week and divide them into no less than 20,000 image frames. Calibrate manually, draw the position of the target in the image, and give the category to which the target belongs. (b) Train the proposed cross-level fusion SSD detection framework. The objective function is obtained by weighting the prior box confidence loss and position loss. The specific form is: The Adam optimizer is used for training. When the training process is fully converged, the detection framework is saved. (c) When deployed in multiple project sites, steps (a)-(b) are repeated according to the specific project, and the detection framework parameters for the specific project site are obtained through fine-tuning.

[0068] In the present invention, the traditional SSD framework is improved, and a shallow feature map is introduced by a cross-level fusion method. In particular, a deep and shallow feature map fusion method based on difference supervision is proposed, thereby enhancing the system's expression ability for small targets and having good detection accuracy. It can detect common loads, occupants and objects in construction elevators, and has good robustness for objects with large scale variations and inevitable occlusion problems in narrow cabin spaces.

[0069] In the video analysis platform proposed by the system, the core algorithm is based on the SSD target detection framework. The traditional SSD detection framework consists of three parts: feature extraction backbone network, dense prior frame generation and prediction. Among them, the specific architecture of the feature extraction backbone network is based on the VGG network, and the deep feature map is subjected to feature fusion operation. However, according to the principle of convolutional neural network, the deep feature map has a large perception area and can better contain the global position information of the target, but the ability to describe small targets is limited. Considering the important role of shallow feature maps in the recognition of small targets, the current improved SSD algorithm is commonly used to directly cascade or add feature maps after aligning the size when fusing deep and shallow feature maps. This type of method actually only describes the first-order statistics of the features, and the expression ability of the obtained model is limited. The present invention proposes to effectively solve the above problems and improve the model's feature extraction ability for small targets by fusing feature maps based on difference supervision. In addition, the original SSD target framework adopts a dense prior frame generation strategy, which is specifically to generate a prior frame point by point on the selected feature map. Under this strategy, the size of the feature map will directly affect the number of generated prior boxes, thereby affecting the number of objects that need to be processed in subsequent classification detection tasks. In order to adapt to the current dense prior box generation method, the size of the fused feature map is aligned downward, that is, the size of the fused feature map is aligned to the size of the deep feature map fused with it through convolution and pooling operations. Figure 1 Avoid introducing too much computational load.

[0070] Considering that construction elevator cars usually use a mesh structure for semi-enclosed processing, the surveillance video in the car is affected by external lighting, the image quality on rainy days and sunny days is quite different, and the color space is not uniform. The detection model relies on image features such as color and texture, so the difference in color space will affect the accuracy of subsequent detection. To solve this problem, the present invention converts the color space of all images and aligns them to the standard value. Specifically, the selection of the standard value is to select the mean of all images in the training set. Implementing this operation in the image preprocessing stage can improve the accuracy of the target detection framework. The specific flow chart is as follows. Figure 5 shown.

[0071] In construction elevators, common loads include passengers, wood, cables, and handheld equipment. The size and shape of the targets vary greatly, including large objects such as wood, mud trucks, etc., and small objects such as handheld electric rotary.

[0072] The present invention first uses an industrial camera to collect video stream images in real time. The camera is installed in the top corner of the industrial elevator farthest from the door opening to achieve the most comprehensive viewing angle. Next, the collected video stream information is transmitted to the back-end server through the video signal forwarding module through the WIFI transmission channel. The server runs a video analysis platform, firstly performs frame processing on the video stream information, extracts image frames by sampling at equal intervals of 7 frames / s, and performs subsequent color space correction, cross-level fusion SSD target detection, anomaly detection, etc. When detecting the target, the feature cross-level fusion method for difference supervision proposed in this paper is used for target detection. In order to enable the detector to learn more actual scenes encountered by industrial elevators in engineering practice, the most widely used face detection benchmark dataset WIDER FICE in the world is used; the engineering hat dataset contains about 10,000 photos of wearing engineering hats in different actual situations; the engineering wooden stick dataset contains about 5,000 photos and the wheelbarrow dataset for pushing mud in engineering contains about 5,000 photos. As described above, the present invention can better meet the requirements of construction elevator target recognition.

Claims

1. Construction elevator standard operation monitoring and alarm method based on video analysis, It is characterized in that A construction elevator standard operation monitoring and alarm system based on video analysis is adopted; the following steps are included: S1. The video analysis platform divides the input video stream into frames, extracts image frames by sampling at equal intervals, and performs a preprocessing operation on the image frames to align the color space to obtain preprocessed image frames; S2. Using a cross-level fusion SSD detection framework to perform target detection on the pre-processed image frames obtained by frame sampling, the target detection environment is a construction elevator; The detection method of the cross-level fusion SSD detection framework in S2 has the following specific steps: (1) Extracting target features In view of the above-mentioned goal, the cross-level fusion SSD algorithm is used to extract features from the training set image frames. The cross-level fusion SSD algorithm not only includes the extraction of deep feature maps by the traditional SSD algorithm, but also adds the extraction of two shallow feature maps on the basis of the traditional SSD algorithm. The sizes of the two newly added shallow feature maps are 150×150 and 76×76. (2) Feature Fusion First, multiple groups of cross-level feature maps are fused based on the two-stream convolutional model. The two cross-level feature maps to be fused in each group are recorded as shallow feature maps f shallow And the deep feature map f deep , the shallow feature map f shallow are the two shallow feature maps in step (1), and the deep feature map f deep is the deep feature map in step (1); where f shallow The size of f is larger, carrying more local features, deep The size of is smaller and carries more global features; the fusion steps are specifically as follows: first, the deep feature map f is deconvolved by a deconvolution operation deep The size of the shallow feature map f shallow Consistent; then calculate the deep feature map f deep Channel outer product B(f shallow, TransConv(f deep ))=f shallow T T ransConv(f deep ) obtains the result f of the bilinear mapping; compresses the dimension of the fusion feature of f through multi-level dense convolution and pooling operations based on difference supervision to obtain a fusion feature map of appropriate size; the method is to use a three-stage series convolution pooling module to compress the dimension of the result of the channel outer product; use the difference supervision method to deconvolve the output of the high-level convolution pooling module, align the size to the adjacent low-level convolution pooling module, calculate the difference between the outputs of the two-level modules, and then perform a hole convolution on the difference to align the size to the fusion feature map to be output; perform inter-level difference supervision on the three-stage series convolution pooling modules respectively, and output the fusion feature; (3) Dense prior box generation The fused feature map obtained in step (2) is used to generate the prior frame. The six sizes are: 38×38, 19×19, 10×10, 5×5, 3×3 and 1×1. Each n×n feature map has n×n center points, and each center point generates k prior frames. The k generated by each center point of the feature map of each layer in the six layers represented by the above six sizes are 4, 6, 6, 6, 4, and 4 respectively. (4) Obtain the type and number of targets in the image Classify and detect the priori frame obtained in step (3) to obtain the type and number of targets in the current input image; each priori frame will be used to perform the following tasks: classify the priori frame through the softmax function to determine whether the priori frame contains a target. If so, first determine the type of target and then determine the number of targets of different types; adjust the size and position of the priori frame through regression analysis to improve the accuracy and efficiency of target classification and detection; (5) Training For the proposed cross-level fusion SSD detection framework, input the training set image frames and loop through steps (1)-(4) to train the model until the target detection accuracy reaches the accuracy required by the actual project, and then the loop is terminated. The trained cross-level fusion SSD detection model is output. S3. Determine whether abnormal behavior that violates safety regulations occurs according to the judgment conditions; The judgment conditions are: the object detected based on the target of step S2 does not belong to the items allowed to enter the construction elevator; the detected passengers exceed the allowed number; if the above conditions are met, it is determined that a violation has occurred, the alarm interface is called, and an alarm signal is output, otherwise it returns to step S1 for the next detection.

2. According to claim 1, the construction elevator standard operation monitoring and alarm method based on video analysis, It is characterized in that The specific method of S1 is as follows: take some preprocessed image frames as training set image frames, first convert the training set image frames from RGB image space to Lab space, and calculate the mean of the three channels of the training set image frames in Lab space respectively. and variance The m L , m a and m b Refers to the mean pixel value of L channel, a channel and b channel respectively, var L , var a and var b They refer to the pixel value variance of L channel, a channel and b channel respectively; Then, the training set image frames are calculated pixel by pixel according to the three channels of L, a and b i∈{L, a, b}, where p i,x ,m i,x and var i,x Respectively represent the pixel value, mean and variance of the pixel x in the current image in channel i; is the adjusted pixel value of pixel point x in channel i; Finally, the image composed of the adjusted pixel values ​​is inversely transformed from the Lab space back to the RGB space to obtain the preprocessed image frame.

3. According to claim 1, the construction elevator standard operation monitoring and alarm method based on video analysis, It is characterized in that In step (2), the specific feature map fusion is: fusing the 150×150 shallow feature map with the 19×19 deep feature map, and fusing the 76×76 shallow feature map with the 38×38 deep feature map.

4. The construction elevator standard operation monitoring and alarm method based on video analysis according to claim 1, It is characterized in that Step (5) includes the following sub-steps: (a) Construction of a monitoring video database of the construction elevator on site: First, a monitoring database based on the project is constructed, which mainly includes the on-site video data captured by the monitoring cameras installed in the construction elevator. The data size is more than 10G, and the number of various targets included should be as balanced as possible. The specific method is to continuously collect elevator monitoring videos for about one week and divide them into no less than 20,000 image frames; calibrate them manually, draw the position of the target in the image and give the category to which the target belongs; (b) The proposed cross-level fusion SSD detection framework is trained. The objective function is obtained by weighting the prior box confidence loss and position loss. The specific form is: Trained using the Adam optimizer; When the training process is fully converged, save the detection framework; (c) When deploying at multiple project sites, repeat steps (a)-(b) according to the specific project, and train the detection framework parameters for the specific project site through fine-tuning.

5. Construction elevator standard operation monitoring and alarm system based on video analysis, It is characterized in that Used to implement the construction elevator standard operation monitoring and alarm method based on video analysis as described in claim 1; the system includes a network camera arranged in the construction elevator car, a video stream transmission module and server, a video analysis platform and an alarm interface; The network camera in the construction elevator car is a collection device for elevator monitoring video; The video stream transmission module, mainly composed of a WIFI communication module, transmits the data collected by the network camera to the server; Server, located at the project site or in the cloud, is mainly used to run the video analysis platform and provide an alarm interface; The video analysis platform mainly performs real-time analysis on the collected video streams and automatically identifies people and typical objects in the shooting field of view.

Citation Information

Patent Citations

  • Intelligent early warning system for construction hoister

    CN202414916U

  • Perimeter target object detecting method and system

    CN109671236A

  • Image target detection method based on FCE-SSD method

    CN113283428A