Method, device and equipment for controlling smoking behavior in logistics area and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-27
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]传统的监管方法为通过物流区域内的监控摄像头获取图像数据,后台监管人员通过监控视频查看图像数据判定物流区域内是否存在抽烟行为,但是,该监管方法需要后台监管人员长期监控多个监控视频,任务繁重
[0021]本发明提供的技术方案中,通过抽烟行为识别模型和动态姿势评估模型,结合每帧监控图像数据的行为识别结果和多帧监控图像数据的动态姿势,结合判断目标人员的抽烟行为自动检测物流区域内的抽烟行为并进行预警,降低后台监管人员的工作量,降低人工成本,提高识别精度,避免误判和漏判,提供多样的抽烟行为管控方案,提高物流安全监管工作效率。
Smart Images

Figure CN115565138B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of logistics area management technology, and in particular to a method, device, equipment and storage medium for controlling smoking behavior in logistics areas. Background Technology
[0002] In the logistics industry, some logistics areas, such as storage areas, inbound / outbound platforms, and integrated control rooms, contain a large number of flammable materials (such as packaging cartons, wooden pallets, etc.), goods, and important control equipment. Smoking in these logistics areas can easily cause fires or even explosions. In order to ensure the safety of property and personnel and the smooth operation of logistics work, it is necessary to supervise and control smoking behavior in logistics areas.
[0003] Traditional monitoring methods rely on surveillance cameras within logistics areas to capture image data. Back-office supervisors then review these images to determine if smoking is occurring. However, this method requires long-term monitoring of multiple video feeds, a demanding task. Existing smoking detection methods rely on single-frame image recognition. Because cigarette butts occupy a small portion of the image, the accuracy is low, leading to false positives and false negatives. This results in a simplistic approach to controlling smoking in logistics areas and low monitoring efficiency. Summary of the Invention
[0004] This invention provides a method, device, equipment, and storage medium for controlling smoking behavior in logistics areas. It is used to automatically detect smoking behavior in logistics areas and issue early warnings, thereby reducing the workload of back-end supervisors, reducing labor costs, improving identification accuracy, avoiding misjudgments and omissions, providing diverse smoking behavior control solutions, and improving the efficiency of logistics safety supervision.
[0005] The first aspect of this invention provides a method for controlling smoking behavior in a logistics area, comprising: acquiring real-time monitoring video data and the logistics area where the monitoring camera is located; performing frame extraction and positioning based on the real-time monitoring video data and a preset target tracking algorithm to obtain multiple monitoring images and the position coordinates of personnel in each monitoring image; inputting the multiple monitoring images and the position coordinates of personnel in each monitoring image into a smoking behavior recognition model to obtain multiple feature maps, and confirming the behavior recognition result of each monitoring image; when the behavior recognition result meets preset conditions, inputting the corresponding multiple feature maps into a preset dynamic posture evaluation model to obtain a smoking behavior evaluation value of the target personnel; when the smoking behavior evaluation value is greater than a first preset value, acquiring the identity information of the target personnel; and generating an early warning processing scheme based on the identity information, the logistics area, and preset access conditions.
[0006] In one feasible implementation, before acquiring real-time monitoring video data and the logistics area of the monitoring camera, the method further includes: acquiring historical monitoring video data from the monitoring camera; and inputting the historical monitoring video data into a preset initial neural network model for training to obtain a smoking behavior recognition model.
[0007] In one feasible implementation, historical surveillance video data is input into a pre-set initial neural network model for training to obtain a smoking behavior recognition model. This includes: preprocessing the historical surveillance video data to obtain a first training set and a second training set; labeling the first training set to obtain feature annotation information; inputting the first training set into the pre-set initial neural network model to obtain a first recognition result; calculating the model's focus loss function based on the first recognition result and the feature annotation information; optimizing the initial neural network model based on the focus loss function to obtain a first neural network model; and inputting the second training set into the first neural network model for training to obtain the smoking behavior recognition model.
[0008] In one feasible implementation, multiple frames of surveillance images and the position coordinates of the person in each frame are input into a smoking behavior recognition model to obtain multiple feature maps, and the behavior recognition result of each frame of surveillance images is confirmed. This includes: acquiring multiple frames of surveillance images, which include at least a first frame and a second frame, wherein the first and second frames satisfy a preset frame extraction time interval; inputting the first frame of surveillance images into the smoking behavior recognition model to obtain multiple feature maps, which include multiple feature maps of different scales, each scale feature map being marked with a target bounding box and a corresponding feature label; obtaining the behavior recognition result of the first frame of surveillance images based on the target bounding box and the corresponding feature label in each scale feature map; and when the behavior recognition result of the first frame of surveillance images meets preset conditions, inputting the second frame of surveillance images into the smoking behavior recognition model based on the position coordinates of the person in each frame of surveillance images to obtain the behavior recognition result of the target person in the second frame of surveillance images.
[0009] In one feasible implementation, when the behavior recognition result meets preset conditions, the corresponding multiple feature maps are input into a preset dynamic posture evaluation model to obtain the smoking behavior evaluation value of the target person. This includes: when the behavior recognition result of the first frame monitoring image and the behavior recognition result of the second frame monitoring image meet preset conditions, according to a preset first scale, obtaining the first feature map and first confidence level corresponding to the first frame monitoring image, and the second feature map and second confidence level corresponding to the second frame monitoring image; determining the first feature map and the second feature map as the video to be evaluated; and inputting the video to be evaluated, the first confidence level, and the second confidence level into the preset dynamic posture evaluation model to obtain the smoking behavior evaluation value of the target person.
[0010] In one feasible implementation, when the smoking behavior assessment value is greater than a first preset value, the identification information of the target person is obtained, including: when the smoking behavior assessment value is greater than the first preset value, according to a preset second scale, the third feature map corresponding to the second frame monitoring image is obtained, the third feature map is a feature map of the target box labeled with the target person's facial tag; based on the third feature map and a preset personnel information database, the identification information of the target person is generated.
[0011] In one feasible implementation, an early warning processing plan is generated based on identity information, logistics area, and preset access conditions, including: obtaining identity information, which may include staff information, visitor information, or stranger information; generating penalty notice information based on staff information and logistics area and sending it to the department where the target personnel are located; generating security reminder information based on visitor information, logistics area, and preset access conditions and sending it to the target personnel and the accessed object; and generating alarm information based on stranger information and logistics area and sending it to the security department.
[0012] A second aspect of the present invention provides a smoking behavior control device for a logistics area, comprising: a first acquisition module for acquiring real-time monitoring video data and the logistics area where the monitoring camera is located; a processing module for performing frame extraction and positioning based on the real-time monitoring video data and a preset target tracking algorithm to obtain multiple monitoring images and the position coordinates of personnel in each monitoring image; a recognition module for inputting the multiple monitoring images and the position coordinates of personnel in each monitoring image into a smoking behavior recognition model to obtain multiple feature maps and confirming the behavior recognition result of each monitoring image; an evaluation module for inputting the corresponding multiple feature maps into a preset dynamic posture evaluation model to obtain a smoking behavior evaluation value of the target personnel when the behavior recognition result meets preset conditions; an identity recognition module for acquiring the identity recognition information of the target personnel when the smoking behavior evaluation value is greater than a first preset value; and a scheme generation module for generating an early warning processing scheme based on the identity recognition information, the logistics area, and preset access conditions.
[0013] In one feasible implementation, a smoking behavior control device for a logistics area further includes: a second acquisition module for acquiring historical monitoring video data from a surveillance camera; and a training module for inputting the historical monitoring video data into a preset initial neural network model for training to obtain a smoking behavior recognition model.
[0014] In one feasible implementation, the training module is specifically used to preprocess historical surveillance video data to obtain a first training set and a second training set; to annotate the first training set to obtain feature annotation information; to input the first training set into a preset initial neural network model to obtain a first recognition result; to calculate the focus loss function of the model based on the first recognition result and the feature annotation information; to optimize the initial neural network model based on the focus loss function to obtain a first neural network model; and to input the second training set into the first neural network model for training to obtain a smoking behavior recognition model.
[0015] In one feasible implementation, the recognition module is specifically used to acquire multiple frames of surveillance images, including at least a first frame and a second frame, wherein the first and second frames satisfy a preset frame extraction time interval; the first frame is input into a smoking behavior recognition model to obtain multiple feature maps, which include feature maps of different scales, each scale feature map being marked with a target bounding box and a corresponding feature label; based on the target bounding box and corresponding feature label in each scale feature map, the behavior recognition result of the first frame is obtained; when the behavior recognition result of the first frame meets preset conditions, the second frame is input into the smoking behavior recognition model based on the position coordinates of the person in each frame of surveillance images to obtain the behavior recognition result of the target person in the second frame.
[0016] In one feasible implementation, the evaluation module is specifically used to, when the behavior recognition results of the first frame monitoring image and the second frame monitoring image meet preset conditions, acquire the first feature map and first confidence level corresponding to the first frame monitoring image, and the second feature map and second confidence level corresponding to the second frame monitoring image according to a preset first scale; determine the first feature map and the second feature map as the video to be evaluated; input the video to be evaluated, the first confidence level and the second confidence level into a preset dynamic posture evaluation model to obtain the smoking behavior evaluation value of the target person.
[0017] In one feasible implementation, the identity recognition module is specifically used to obtain a third feature map corresponding to the second frame monitoring image according to a preset second scale when the smoking behavior assessment value is greater than a first preset value. The third feature map is a feature map of the target box labeled with the facial tag of the target person. Based on the third feature map and a preset personnel information database, the identity recognition information of the target person is generated.
[0018] In one feasible implementation, the scheme generation module is specifically used to obtain identity recognition information, including staff information, visitor information, or stranger information; generate penalty publicity information based on staff information and logistics area and send it to the department where the target personnel are located; generate security reminder information based on visitor information, logistics area, and preset access conditions and send it to the target personnel and the accessed object; and generate alarm information based on stranger information and logistics area and send it to the security department.
[0019] A third aspect of the present invention provides a smoking behavior control device for a logistics area, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the smoking behavior control device for the logistics area to execute the above-described smoking behavior control method for the logistics area.
[0020] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the above-described method for controlling smoking behavior in a logistics area.
[0021] The technical solution provided by this invention uses a smoking behavior recognition model and a dynamic posture evaluation model. By combining the behavior recognition results of each frame of monitoring image data and the dynamic posture of multiple frames of monitoring image data, the smoking behavior of the target personnel is judged to automatically detect smoking behavior in the logistics area and issue warnings. This reduces the workload of back-end supervisors, lowers labor costs, improves recognition accuracy, avoids misjudgments and omissions, provides diverse smoking behavior control solutions, and improves the efficiency of logistics safety supervision. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of an embodiment of the method for controlling smoking behavior in the logistics area according to the present invention;
[0023] Figure 2 This is a schematic diagram of another embodiment of the method for controlling smoking behavior in the logistics area according to the present invention;
[0024] Figure 3 This is a schematic diagram of an embodiment of the smoking control device in the logistics area according to the present invention;
[0025] Figure 4 This is a schematic diagram of another embodiment of the smoking control device in the logistics area according to the present invention;
[0026] Figure 5 This is a schematic diagram of an embodiment of the smoking control device in the logistics area according to an embodiment of the present invention. Detailed Implementation
[0027] This invention provides a method, device, equipment, and storage medium for controlling smoking behavior in logistics areas. It is used to automatically detect smoking behavior in logistics areas and issue early warnings, thereby reducing the workload of back-end supervisors, reducing labor costs, improving identification accuracy, avoiding misjudgments and omissions, providing diverse smoking behavior control solutions, and improving the efficiency of logistics safety supervision.
[0028] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” or “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 One embodiment of the method for controlling smoking behavior in logistics areas according to the present invention includes:
[0030] 101. Obtain real-time monitoring video data and the logistics area where the monitoring cameras are located.
[0031] It is understood that the implementing entity of this invention can be a smoking behavior control device in a logistics area, a monitoring center control terminal in a logistics park, or a server; the specific implementation is not limited here. This embodiment of the invention will be described using a monitoring center control terminal as an example.
[0032] The monitoring center control terminal acquires surveillance video data from all logistics areas within the logistics park. This video data is collected from multiple surveillance cameras within the logistics park, as well as surveillance cameras at the access control points of each logistics area. Each surveillance camera has a unique device identification code. The surveillance cameras transmit the collected video data and the corresponding device identification code to the monitoring center control terminal in real time via wired or wireless means. The monitoring center control terminal stores a location relationship table, and the logistics area where the surveillance camera is installed can be determined by querying the location relationship table using the device identification code.
[0033] The logistics area in this embodiment includes warehouses, outbound platforms, inbound platforms, a comprehensive control room, equipment areas, office areas, parking lots, and other logistics work areas. Warehouses can be divided into picking areas, shipping areas, storage areas, and transit areas, or can be zoned according to actual conditions. Each logistics area is assigned a corresponding no-smoking safety level, which includes at least Level 1, Level 2, and Level 3. Level 1 is a complete no-smoking zone, Level 2 is a partial no-smoking zone, and Level 3 is no-smoking zone. Generally, warehouses and equipment areas are Level 1, but the specific no-smoking safety level can be set for each logistics area based on actual circumstances.
[0034] 102. Based on real-time monitoring video data and a preset target tracking algorithm, perform frame extraction and positioning to obtain multiple monitoring images and the position coordinates of personnel in each monitoring image.
[0035] The monitoring center control terminal extracts frames from real-time monitoring video data at preset frame extraction intervals to obtain multiple monitoring images. Based on the multiple monitoring images and a preset target tracking algorithm, personnel in each monitoring image can be located, and their position coordinates can be obtained. When a person is identified as smoking in the first monitoring image, that person is identified as the target person. By using the position coordinates of the person in each monitoring image, the position of the target person in subsequent monitoring images can be confirmed. Then, feature extraction and enhancement are performed only on that location area, while ignoring other people and the background in subsequent monitoring images, which can reduce the computational cost of the smoking behavior recognition model.
[0036] In this embodiment, the frame extraction time interval includes a first time interval and a second time interval. The first time interval can be set according to the computing resources of the monitoring center control terminal. When computing resources are sufficient, it can be omitted or a shorter frame extraction time interval can achieve better monitoring of smoking behavior. When computing resources are insufficient, a longer frame extraction time interval can be set to reduce the consumption of computing resources. When the confidence level of smoking behavior in the behavior recognition result obtained by the smoking behavior recognition model of the first frame monitoring image is greater than the preset confidence level, frames are extracted according to the second time interval to obtain subsequent frame monitoring images. The number of frames in the subsequent frame monitoring images is related to the second time interval. The subsequent frame monitoring images include at least the second frame monitoring image and the third frame monitoring image. The second time interval can be confirmed by the smoking duration of a general person. By setting an appropriate second time interval, the first frame monitoring image and the subsequent frame monitoring images that meet the requirements are merged into a video to be evaluated. This video can be used to evaluate the dynamic posture of the target person over a certain period of time, such as changes in gestures, thereby further improving the accuracy of smoking behavior judgment.
[0037] 103. Input multiple frames of surveillance images and the location coordinates of the person in each frame of surveillance images into the smoking behavior recognition model to obtain multiple feature maps, and confirm the behavior recognition results of each frame of surveillance images.
[0038] The monitoring center control terminal acquires multiple frames of monitoring images, including at least a first frame, a second frame, and a third frame. The first frame is input into the smoking behavior recognition model for processing to obtain the first recognition result of the first frame. If the first recognition result does not meet the preset conditions, it is determined that there is no smoking behavior, and recognition continues.
[0039] When the first recognition result meets the preset conditions, the location of the target person in the second frame of the monitoring image is obtained based on the position coordinates of the person in each frame of the monitoring image. The attention weight of that location is increased, and the second frame of the monitoring image is input into the smoking behavior recognition model to obtain the second recognition result of the target person. When the second recognition result does not meet the preset conditions, the second time interval is adjusted and frames are re-sampled, or it is determined whether the person is tracked incorrectly. If so, the attention weight of that location is deleted and re-recognition is performed. When the re-recognition result still does not meet the preset conditions, it is determined that there is no smoking behavior and the first frame of the monitoring image is identified incorrectly. The parameters of the smoking behavior recognition model are adjusted according to the focus loss function of the first frame of the monitoring image.
[0040] When the second recognition result meets the preset conditions, the third frame of the monitoring image is input to obtain the third recognition result; when the third recognition result does not meet the preset conditions, the above steps of adjusting the second time interval, re-recognizing, and optimizing parameters are repeated; when the third recognition result meets the preset conditions, recognition is stopped.
[0041] It should be further explained that the first, second, and third monitoring images mentioned above are only examples. In practice, a second time interval can be set to acquire multiple monitoring images for identification. The more frames identified and the shorter the second time interval, the more accurate the judgment result will be. Of course, the greater the computing resources consumed, the more frames and the shorter the second time interval should be considered to obtain a more accurate identification result.
[0042] The smoking behavior recognition model used in this application is an improved RetinaNet network model. The original RetinaNet network model includes a backbone feature extraction network ResNet layer, a feature pyramid network (FPN) layer, and a regression network layer. The regression network layer includes a box subnet and a class subnet. The improved RetinaNet network model introduces an attention (Squeeze-and-Excitation, SE) layer. By introducing a channel attention mechanism to enhance the feature map, it can further utilize the correlation between extracted channels to enhance effective features and suppress ineffective features.
[0043] Taking the first frame of a surveillance image as an example, the operation of the improved RetinaNet network model is explained as follows: The first frame of the surveillance image is input into the backbone feature extraction network of the smoking behavior recognition model to label the backbone features of the person, thus obtaining a backbone feature map; the backbone feature map is input into the attention layer of the smoking behavior recognition model to label the key attention regions, thus obtaining an attention feature map; the attention feature map is input into the feature pyramid network layer of the smoking behavior recognition model, and residual connections and feature fusion are performed according to the preset scale gradient to obtain multiple scale feature maps; the multiple scale feature maps are input into the regression network layer of the smoking behavior recognition model to perform target box labeling and classification regression for each scale feature map, thus obtaining the corresponding regression feature map. The presence or absence of smoking behavior can be determined by the regression features.
[0044] The backbone feature map, due to multiple convolutions, may lose features of small objects, which is not conducive to the detection of small objects, but is beneficial for the detection of large objects. It processes the image input from multiple channels of the backbone feature extraction network through the attention layer, assigning corresponding attention weights to each channel. This can remove noise information that is irrelevant to the recognition task. In the recognition of the second frame of the surveillance image, the attention layer can increase the attention weight of the location area according to the location coordinates of the target person, so as to improve the recognition speed of the model. This is also applicable to other subsequent frames of surveillance images. Furthermore, the attention feature maps can be fused with multi-scale features through the feature pyramid network layer. The feature pyramid conveys strong localization features from the bottom up, aggregating parameters from different backbone layers to different detection layers, and outputting multiple scale feature maps for prediction. This allows for the extraction of higher semantic information and takes into account the detection of both large and small objects. Through the feature pyramid network layer, five scale feature maps can be obtained, or other numbers. Each scale feature map has the same image size but a different region of interest. The scale feature map output at the bottom of the feature pyramid can be the first layer image containing the entire body of a person, followed by the second layer image containing the upper body, the third layer image containing the face, the fourth layer image containing the hands, and the fifth layer image containing the mouth. These five scale feature maps are then fed into the regression network layer, and the size of the bounding boxes is set according to the corresponding scale. The bounding boxes are then labeled to obtain the bounding box regions corresponding to the five scale feature maps. Next, target bounding box classification regression is performed to obtain feature labels for target bounding boxes at five scales. The feature labels include human image labels and target labels. Target labels include at least cigarettes, cigarette butts, and smoke. Human image labels can include the whole body, face, hands, mouth, etc. Based on the overlap of feature labels and target bounding box regions, or judging whether the relative distance of the target bounding boxes is less than a threshold, the first recognition result of the first frame of the surveillance image is obtained. The first recognition result can directly represent whether smoking behavior exists or not, or it can be represented as the confidence level of the existence of smoking behavior.
[0045] 104. When the behavior recognition result meets the preset conditions, the corresponding multiple feature maps are input into the preset dynamic posture evaluation model to obtain the smoking behavior evaluation value of the target person.
[0046] When the behavior recognition results of the first frame, second frame, and third frame monitoring images determined by the monitoring center control terminal in step 103 above all meet the preset conditions, the first feature map and first confidence level corresponding to the first frame monitoring image, the second feature map and second confidence level corresponding to the second frame monitoring image, and the third feature map and third confidence level corresponding to the third frame monitoring image are obtained according to the preset first scale. The first feature map, second feature map, and third feature map are merged into the video to be evaluated. The video to be evaluated, the first confidence level, the second confidence level, and the third confidence level are input into the preset dynamic posture evaluation model to obtain the evaluation value of the smoking behavior of the target person by combining the static action evaluation result (i.e., confidence level) of each frame monitoring image data and the dynamic action evaluation in the video to be evaluated.
[0047] In this embodiment, the preset first scale can be a fourth layer image that includes a person's hand. By judging the trajectory or movement changes of the hand posture in multiple monitoring images, it can be determined whether there is smoking behavior.
[0048] It should be further explained that the monitoring center control terminal can also set different scales for the first frame, the second frame, and the third frame of the monitoring image, and respectively acquire the second layer image containing the upper body area of the person in the first frame, the third layer image containing the face of the person in the second frame, and the fourth layer image containing the hands of the person in the third frame. These images are then merged into a video to be evaluated, and the smoking behavior is evaluated based on the motion trajectory of the target box carrying the target label in the video to be evaluated.
[0049] 105. When the smoking behavior assessment value is greater than the first preset value, obtain the target person's identity information.
[0050] When the smoking behavior assessment value exceeds a first preset value, the monitoring center control terminal acquires a third-layer image containing a human face from the feature map output by the smoking recognition model. Based on the third-layer image and the facial recognition algorithm, it performs recognition and matching on the facial information stored in the monitoring center control terminal to obtain the target person's identity information. When the match is successful, the identity type is determined to be either staff or visitor, and the corresponding identity information can be obtained. When the identity type is staff, the identity information may include the identity type, department, accessible logistics areas, and contact information of the department head. When the identity type is visitor, the identity information may include the identity type, permitted logistics areas, the visitor's contact mobile phone number, and information about the person being visited. When the match fails, the identity type is determined to be stranger, and the identity information consists of the identity type and the third-layer image containing a human face.
[0051] Visitors are individuals who are not staff members of the logistics park, but whose facial information was captured and whose visit information was registered by the access control surveillance camera when they entered the logistics park. The visit information includes at least the logistics area visited, the purpose of the visit, the contact mobile phone number, and the person being visited. The person being visited is a staff member of the logistics park.
[0052] 106. Generate an early warning processing plan based on identity information, logistics area, and preset access conditions.
[0053] Based on the target person's identity information, the logistics area where the target person smoked, and preset access conditions, the monitoring center control terminal generates early warning and handling schemes of different security levels. For example, when the identity type is a stranger, and the logistics area where the target person smoked has a corresponding no-smoking security level of Level 1 (i.e., no smoking in the entire area), the highest level early warning and handling scheme is generated, triggering the alarm bell in the area and sending a third-layer image containing the person's face to the security department.
[0054] The access conditions in this embodiment may include that the identity type is not a stranger and that the person's smoking behavior assessment value is consistent with the no-smoking safety level of the logistics area. If any condition is not met, it is determined that the access conditions are not met.
[0055] The technical solution provided by this invention uses a smoking behavior recognition model and a dynamic posture evaluation model. By combining the behavior recognition results of each frame of monitoring image data and the dynamic posture of multiple frames of monitoring image data, the smoking behavior of the target personnel is judged to automatically detect smoking behavior in the logistics area and issue warnings. This reduces the workload of back-end supervisors, lowers labor costs, improves recognition accuracy, avoids misjudgments and omissions, provides diverse smoking behavior control solutions, and improves the efficiency of logistics safety supervision.
[0056] Please see Figure 2 Another embodiment of the method for controlling smoking behavior in logistics areas according to the present invention includes:
[0057] 201. Obtain historical surveillance video data from the surveillance cameras.
[0058] The monitoring center control terminal acquires historical video data from surveillance cameras. It can filter the acquired historical video data, balance the ratio between samples showing smoking behavior and those not showing smoking behavior, and perform image processing such as deblurring, illumination distortion correction, cropping, horizontal flipping, and rotation on the extracted historical surveillance images to obtain the cleanest possible historical surveillance images.
[0059] 202. Input historical surveillance video data into a preset initial neural network model for training to obtain a smoking behavior recognition model.
[0060] The monitoring center control terminal preprocesses historical surveillance video data, dividing it into a first training set and a second training set. The first training set is labeled to obtain feature annotation information. This first training set is then input into a pre-set initial neural network model to obtain a third recognition result. Based on the third recognition result and the feature annotation information, the model's focus loss function is calculated. The initial neural network model is then optimized based on the focus loss function to obtain a first neural network model. The second training set is then input into the first neural network model for training, resulting in a smoking behavior recognition model. The first training set contains fewer samples than the second training set. The first round of model training using the first training set freezes certain parameters (such as the number of samples), allowing for rapid model optimization with a small sample size. The second round of model training using the second training set unfreezes the parameters and allows for final model optimization with a larger sample size. Whether optimization is complete can be determined by the convergence of the focus function.
[0061] In one feasible implementation, the first training set is labeled to obtain feature annotation information, including: the monitoring center control terminal saves the first training set to a file to be labeled according to a preset file format; and annotates the file to be labeled using an annotation tool to obtain feature annotation information. For example, first, open the labelimg.exe file. Create a folder named JPEGImages to store the monitoring images to be labeled, and place the labels annotated by labelimg in the Annotations folder; open the folder containing the monitoring images; select Auto Saving in the view to automatically save; use bounding boxes to mark the positions and create labels; after the annotation is completed, the Annotations folder will automatically generate a feature annotation information .xml file.
[0062] In one feasible implementation, the monitoring center control terminal acquires a second training set, which includes at least the first and second monitoring image data. The first monitoring image data is input into the backbone feature extraction network layer of the first neural network model to label the backbone features of personnel, obtaining corresponding backbone feature maps. The backbone feature maps are then input into the attention layer of the first neural network model to label key attention regions, obtaining attention feature maps. The attention feature maps are input into the feature pyramid network layer of the first neural network model, where residual connections and feature fusion are performed according to a preset scale gradient, resulting in multiple scale feature maps. These multiple scale feature maps are then input into the regression network layer of the smoking behavior recognition model to perform bounding box labeling and classification regression for each scale feature map, obtaining a fourth recognition result. A focus loss function is calculated based on the fourth recognition result. The parameters of the first neural network model are updated according to the focus loss function, and the second monitoring image data is input for training. When the focus loss function meets a preset convergence condition, parameter updates are stopped, resulting in the smoking recognition model.
[0063] In one feasible implementation, to address the convergence difficulties caused by class imbalance in smoking recognition models, the model parameters are optimized by calculating a focus loss function. Class imbalance refers to the ratio of positive / negative samples to easy / difficult samples. Positive / negative samples are surveillance image samples containing smoking behavior and other surveillance image samples not containing smoking behavior. Easy / difficult samples include easy and difficult samples containing smoking behavior (i.e., although smoking behavior exists, the smoking recognition model struggles to identify the behavior due to factors such as the target person's posture). The focus function includes positive / negative sample coefficients and easy / difficult sample coefficients, which influence each other. When evaluating accuracy, both need to be adjusted together. The loss value calculated by the focus function includes two parts: the target box regression loss value and the classification loss value, used to determine the difference between the model's labeled target box and the correct classification value.
[0064] 203. Obtain real-time monitoring video data and the logistics area where the monitoring cameras are located.
[0065] 204. Based on real-time monitoring video data and a preset target tracking algorithm, perform frame extraction and positioning to obtain multiple monitoring images and the position coordinates of personnel in each monitoring image.
[0066] Steps 203-204 are similar to steps 101-102 above, and will not be repeated here.
[0067] 205. Input multiple frames of surveillance images and the location coordinates of the person in each frame of surveillance images into the smoking behavior recognition model to obtain multiple feature maps, and confirm the behavior recognition results of each frame of surveillance images.
[0068] The monitoring center control terminal acquires multiple frames of monitoring images, including at least a first frame and a second frame, with the first and second frames meeting a preset frame extraction time interval. The first frame is input into a smoking behavior recognition model to obtain multiple feature maps, including feature maps of different scales. Each feature map at each scale is marked with a target bounding box and a corresponding feature label. Based on the target bounding box and corresponding feature label in each feature map at each scale, the behavior recognition result of the first frame is obtained. When the behavior recognition result of the first frame meets preset conditions, the second frame is input into the smoking behavior recognition model based on the position coordinates of the person in each frame to obtain the behavior recognition result of the target person in the second frame.
[0069] It should be further explained that the multi-frame monitoring image may also include the third, fourth, up to the Nth frame monitoring image, and the number of monitoring image frames to be identified is determined according to the actual situation.
[0070] 206. When the behavior recognition result meets the preset conditions, the corresponding multiple feature maps are input into the preset dynamic posture evaluation model to obtain the smoking behavior evaluation value of the target person.
[0071] When the behavior recognition results of the first frame monitoring image and the second frame monitoring image meet the preset conditions, the monitoring center control terminal acquires the first feature map and first confidence level corresponding to the first frame monitoring image, and the second feature map and second confidence level corresponding to the second frame monitoring image, according to the preset first scale; the first feature map and the second feature map are determined as the video to be evaluated; the video to be evaluated, the first confidence level and the second confidence level are input into the preset dynamic posture evaluation model to obtain the evaluation value of the smoking behavior of the target person.
[0072] 207. When the smoking behavior assessment value is greater than the first preset value, obtain the target person's identity information.
[0073] When the smoking behavior assessment value is greater than the first preset value, the monitoring center control terminal acquires the third feature map corresponding to the second frame monitoring image according to the preset second scale. The third feature map is a feature map with the target person's facial label marked on the target box. Based on the third feature map and the preset personnel information database, the identity recognition information of the target person is generated.
[0074] 208. Generate an early warning processing plan based on identity information, logistics area, and preset access conditions.
[0075] The monitoring center control terminal acquires identification information, including staff information, visitor information, or stranger information; based on staff information and logistics area, it generates penalty notice information and sends it to the department where the target personnel are located; based on visitor information, logistics area, and preset access conditions, it generates security reminder information and sends it to the target personnel and the accessed object; based on stranger information and logistics area, it generates alarm information and sends it to the security department.
[0076] In this embodiment of the invention, the initial neural network model is trained twice, and the model is optimized through a focus function. This solves the difficulties caused by positive and negative samples and easy and difficult samples in optimizing the smoking behavior recognition model. By combining the smoking behavior recognition model and the dynamic posture evaluation model, and combining the behavior recognition results of each frame of monitoring image data with the dynamic posture of multiple frames of monitoring image data, the smoking behavior of the target personnel is judged to automatically detect smoking behavior in the logistics area and issue warnings. This reduces the workload of back-end supervisors, reduces labor costs, improves recognition accuracy, avoids misjudgments and omissions, provides diverse smoking behavior control solutions, and improves the efficiency of logistics safety supervision.
[0077] The above describes the method for controlling smoking behavior in the logistics area according to embodiments of the present invention. The following describes the device for controlling smoking behavior in the logistics area according to embodiments of the present invention. Please refer to [link / reference]. Figure 3 One embodiment of the smoking control device in the logistics area of the present invention includes:
[0078] The first acquisition module 301 is used to acquire real-time monitoring video data and the logistics area where the monitoring camera is located.
[0079] Processing module 302 is used to extract frames and locate objects based on real-time monitoring video data and a preset target tracking algorithm to obtain multiple monitoring images and the position coordinates of people in each monitoring image.
[0080] The recognition module 303 is used to input multiple frames of monitoring images and the position coordinates of the person in each frame of monitoring images into the smoking behavior recognition model to obtain multiple feature maps and confirm the behavior recognition result of each frame of monitoring images.
[0081] The evaluation module 304 is used to input multiple corresponding feature maps into a preset dynamic posture evaluation model when the behavior recognition result meets the preset conditions, so as to obtain the smoking behavior evaluation value of the target person.
[0082] The identity recognition module 305 is used to obtain the identity recognition information of the target person when the smoking behavior assessment value is greater than the first preset value;
[0083] The scheme generation module 306 is used to generate an early warning processing scheme based on identity information, logistics area and preset access conditions.
[0084] In this embodiment of the invention, a smoking behavior recognition model and a dynamic posture evaluation model are used. By combining the behavior recognition results of each frame of monitoring image data and the dynamic posture of multiple frames of monitoring image data, the smoking behavior of the target personnel is judged to automatically detect smoking behavior in the logistics area and issue an early warning. This reduces the workload of back-end supervisors, lowers labor costs, improves recognition accuracy, avoids misjudgment and missed judgment, provides diverse smoking behavior control solutions, and improves the efficiency of logistics safety supervision.
[0085] Please see Figure 4 Another embodiment of the smoking control device in the logistics area of this invention includes:
[0086] The first acquisition module 301 is used to acquire real-time monitoring video data and the logistics area where the monitoring camera is located.
[0087] Processing module 302 is used to extract frames and locate objects based on real-time monitoring video data and a preset target tracking algorithm to obtain multiple monitoring images and the position coordinates of people in each monitoring image.
[0088] The recognition module 303 is used to input multiple frames of monitoring images and the position coordinates of the person in each frame of monitoring images into the smoking behavior recognition model to obtain multiple feature maps and confirm the behavior recognition result of each frame of monitoring images.
[0089] The evaluation module 304 is used to input multiple corresponding feature maps into a preset dynamic posture evaluation model when the behavior recognition result meets the preset conditions, so as to obtain the smoking behavior evaluation value of the target person.
[0090] The identity recognition module 305 is used to obtain the identity recognition information of the target person when the smoking behavior assessment value is greater than the first preset value;
[0091] The scheme generation module 306 is used to generate an early warning processing scheme based on identity information, logistics area and preset access conditions.
[0092] Optionally, the smoking control device in the logistics area may also include:
[0093] The second acquisition module 307 is used to acquire historical surveillance video data from the surveillance camera;
[0094] Training module 308 is used to input historical surveillance video data into a preset initial neural network model for training to obtain a smoking behavior recognition model.
[0095] Optionally, the training module 308 is specifically used for: preprocessing historical surveillance video data to obtain a first training set and a second training set; labeling the first training set to obtain feature labeling information; inputting the first training set into a preset initial neural network model to obtain a first recognition result; calculating the focus loss function of the model based on the first recognition result and the feature labeling information; optimizing the initial neural network model based on the focus loss function to obtain a first neural network model; and inputting the second training set into the first neural network model for training to obtain a smoking behavior recognition model.
[0096] Optionally, the recognition module 303 is specifically used to acquire multiple frames of monitoring images, including at least a first frame and a second frame, wherein the first and second frames satisfy a preset frame extraction time interval; input the first frame of monitoring images into the smoking behavior recognition model to obtain multiple feature maps, which include feature maps of different scales, each scale feature map being marked with a target bounding box and a corresponding feature label; based on the target bounding box and corresponding feature label in each scale feature map, the behavior recognition result of the first frame of monitoring images is obtained; when the behavior recognition result of the first frame of monitoring images meets preset conditions, based on the position coordinates of the person in each frame of monitoring images, the second frame of monitoring images is input into the smoking behavior recognition model to obtain the behavior recognition result of the target person in the second frame of monitoring images.
[0097] Optionally, the evaluation module 304 is specifically used to, when the behavior recognition results of the first frame monitoring image and the second frame monitoring image meet the preset conditions, obtain the first feature map and the first confidence level corresponding to the first frame monitoring image, and the second feature map and the second confidence level corresponding to the second frame monitoring image according to the preset first scale; determine the first feature map and the second feature map as the video to be evaluated; input the video to be evaluated, the first confidence level and the second confidence level into the preset dynamic posture evaluation model to obtain the smoking behavior evaluation value of the target person.
[0098] Optionally, the identity recognition module 305 is specifically used to obtain the third feature map corresponding to the second frame monitoring image according to the preset second scale when the smoking behavior assessment value is greater than the first preset value. The third feature map is a feature map of the target box labeled with the facial tag of the target person. Based on the third feature map and the preset personnel information database, the identity recognition information of the target person is generated.
[0099] Optionally, the scheme generation module 306 is specifically used to obtain identity recognition information, including staff information, visitor information, or stranger information; generate penalty publicity information based on staff information and logistics area and send it to the department where the target personnel are located; generate security reminder information based on visitor information, logistics area, and preset access conditions and send it to the target personnel and the accessed object; and generate alarm information based on stranger information and logistics area and send it to the security department.
[0100] In this embodiment of the invention, the initial neural network model is trained twice, and the model is optimized through a focus function. This solves the difficulties caused by positive and negative samples and easy and difficult samples in optimizing the smoking behavior recognition model. By combining the smoking behavior recognition model and the dynamic posture evaluation model, and combining the behavior recognition results of each frame of monitoring image data with the dynamic posture of multiple frames of monitoring image data, the smoking behavior of the target personnel is judged to automatically detect smoking behavior in the logistics area and issue warnings. This reduces the workload of back-end supervisors, reduces labor costs, improves recognition accuracy, avoids misjudgments and omissions, provides diverse smoking behavior control solutions, and improves the efficiency of logistics safety supervision.
[0101] above Figure 3 and Figure 4 The smoking behavior control device in the logistics area of this invention is described in detail from the perspective of modular functional entities. The smoking behavior control device in the logistics area of this invention is described in detail from the perspective of hardware processing.
[0102] Figure 5 This is a schematic diagram of a smoking control device 500 in a logistics area according to an embodiment of the present invention. The smoking control device 500 can vary significantly due to different configurations or performance characteristics. It may include one or more central processing units (CPUs) 510 (e.g., one or more processors) and a memory 520, and one or more storage media 530 (e.g., one or more mass storage devices) for storing application programs 533 or data 532. The memory 520 and storage media 530 can be temporary or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the smoking control device 500 in the logistics area. Furthermore, the processor 510 may be configured to communicate with the storage media 530 and execute the series of instruction operations in the storage media 530 on the smoking control device 500 in the logistics area.
[0103] The smoking control device 500 in the logistics area may also include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input / output interfaces 560, and / or one or more operating systems 531, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 5 The illustrated structure of the smoking control equipment in the logistics area does not constitute a limitation on the smoking control equipment in the logistics area. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0104] The present invention also provides a smoking behavior control device for a logistics area. The computer device includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor performs the steps of the smoking behavior control method for the logistics area in the above embodiments.
[0105] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of a method for controlling smoking behavior in a logistics area.
[0106] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0107] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0108] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for controlling smoking behavior in a logistics area, characterized in that, The methods for controlling smoking in the logistics area include: Obtain real-time monitoring video data and the logistics area where the monitoring cameras are located; Based on real-time monitoring video data and a pre-set target tracking algorithm, frames are extracted and located to obtain multiple monitoring images and the position coordinates of people in each monitoring image. The multiple frames of surveillance images and the location coordinates of the person in each frame of surveillance images are input into the smoking behavior recognition model to obtain multiple feature maps, and the behavior recognition result of each frame of surveillance images is confirmed. When the behavior recognition result meets the preset conditions, the corresponding multiple feature maps are input into the preset dynamic posture evaluation model to obtain the smoking behavior evaluation value of the target person. When the smoking behavior assessment value is greater than a first preset value, the target person's identity information is obtained. Based on the identity information, the logistics area, and the preset access conditions, an early warning processing plan is generated; The process of inputting the multiple frames of surveillance images and the location coordinates of the person in each frame of surveillance images into the smoking behavior recognition model to obtain multiple feature maps and confirming the behavior recognition result of each frame of surveillance images includes: The multi-frame monitoring image is acquired, and the multi-frame monitoring image includes at least a first frame monitoring image and a second frame monitoring image, wherein the first frame monitoring image and the second frame monitoring image satisfy a preset frame extraction time interval; The first frame of the monitoring image is input into the smoking behavior recognition model to obtain multiple feature maps. The multiple feature maps include multiple feature maps of different scales. Each feature map of a scale is marked with a target box and a corresponding feature label. Based on the target bounding box and corresponding feature label marked in the feature map of each scale, the behavior recognition result of the first frame of the surveillance image is obtained; when the corresponding person is identified as smoking in the first frame of the surveillance image, the corresponding person is identified as the target person. When the behavior recognition result of the first frame monitoring image meets the preset conditions, the location of the target person in the second frame monitoring image is obtained according to the position coordinates of the person in each frame monitoring image. The attention weight of the location is increased, and the second frame monitoring image is input into the smoking behavior recognition model to obtain the behavior recognition result of the target person in the second frame monitoring image.
2. The method for controlling smoking behavior in logistics areas according to claim 1, characterized in that, Before acquiring the real-time monitoring video data and the logistics area of the monitoring camera, the method further includes: Obtain historical surveillance video data from surveillance cameras; The historical surveillance video data is input into a preset initial neural network model for training to obtain a smoking behavior recognition model.
3. The method for controlling smoking behavior in logistics areas according to claim 2, characterized in that, The step of inputting the historical surveillance video data into a preset initial neural network model for training to obtain a smoking behavior recognition model includes: The historical surveillance video data is preprocessed to obtain the first training set and the second training set; The first training set is labeled to obtain feature annotation information; The first training set is input into a preset initial neural network model to obtain the first recognition result; Based on the first recognition result and the feature annotation information, calculate the focus loss function of the model; The initial neural network model is optimized based on the focus loss function to obtain the first neural network model; The second training set is input into the first neural network model for training to obtain a smoking behavior recognition model.
4. The method for controlling smoking behavior in logistics areas according to claim 1, characterized in that, When the behavior recognition result meets preset conditions, the corresponding multiple feature maps are input into a preset dynamic posture evaluation model to obtain the smoking behavior evaluation value of the target person, including: When the behavior recognition results of the first frame monitoring image and the behavior recognition results of the second frame monitoring image meet the preset conditions, the first feature map and the first confidence level corresponding to the first frame monitoring image, and the second feature map and the second confidence level corresponding to the second frame monitoring image are obtained according to the preset first scale. The first feature map and the second feature map are identified as the video to be evaluated; The video to be evaluated, the first confidence level, and the second confidence level are input into a preset dynamic posture evaluation model to obtain the evaluation value of the target person's smoking behavior.
5. The method for controlling smoking behavior in logistics areas according to claim 4, characterized in that, When the smoking behavior assessment value is greater than a first preset value, the identification information of the target person is obtained, including: When the smoking behavior assessment value is greater than the first preset value, the third feature map corresponding to the second frame monitoring image is obtained according to the preset second scale. The third feature map is a feature map in which the target box is labeled with the facial tag of the target person. Based on the third feature map and the pre-set personnel information database, the identity information of the target personnel is generated.
6. The method for controlling smoking behavior in logistics areas according to any one of claims 1-5, characterized in that, The step of generating an early warning processing plan based on the identity information, the logistics area, and preset access conditions includes: The identity recognition information is obtained, including staff information, visitor information, or stranger information; Based on the staff information and the logistics area, generate a penalty notice and send it to the department where the target personnel are located; Based on the visitor information, the logistics area, and the preset access conditions, a security alert message is generated and sent to the target personnel and the accessed object; Based on the stranger information and the logistics area, an alarm message is generated and sent to the security department.
7. A device for controlling smoking behavior in a logistics area, characterized in that, The smoking control device in the logistics area includes: The first acquisition module is used to acquire real-time monitoring video data and the logistics area where the monitoring camera is located. The processing module is used to extract frames and locate objects based on real-time monitoring video data and a preset target tracking algorithm, so as to obtain multiple monitoring images and the position coordinates of people in each monitoring image. The recognition module is used to input the multiple frames of monitoring images and the location coordinates of the person in each frame of monitoring images into the smoking behavior recognition model to obtain multiple feature maps and confirm the behavior recognition result of each frame of monitoring images. The evaluation module is used to input multiple corresponding feature maps into a preset dynamic posture evaluation model when the behavior recognition result meets preset conditions, so as to obtain the smoking behavior evaluation value of the target person. The identity recognition module is used to obtain the identity recognition information of the target person when the evaluation value of the smoking behavior is greater than a first preset value; The scheme generation module is used to generate an early warning processing scheme based on the identity recognition information, the logistics area, and preset access conditions; The identification module is also used to acquire the multi-frame monitoring images, which include at least a first frame monitoring image and a second frame monitoring image, wherein the first frame monitoring image and the second frame monitoring image satisfy a preset frame extraction time interval; The first frame of the monitoring image is input into the smoking behavior recognition model to obtain multiple feature maps. The multiple feature maps include multiple feature maps of different scales. Each feature map of a scale is marked with a target box and a corresponding feature label. Based on the target bounding box and corresponding feature label marked in the feature map of each scale, the behavior recognition result of the first frame of the surveillance image is obtained; when the corresponding person is identified as smoking in the first frame of the surveillance image, the corresponding person is identified as the target person. When the behavior recognition result of the first frame monitoring image meets the preset conditions, the location of the target person in the second frame monitoring image is obtained according to the position coordinates of the person in each frame monitoring image. The attention weight of the location is increased, and the second frame monitoring image is input into the smoking behavior recognition model to obtain the behavior recognition result of the target person in the second frame monitoring image.
8. A device for controlling smoking behavior in a logistics area, characterized in that, The smoking control device in the logistics area includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the smoking behavior control device in the logistics area to execute the smoking behavior control method in the logistics area as described in any one of claims 1-6.
9. A computer-readable storage medium storing instructions thereon, characterized in that, When the instruction is read and executed, it performs the smoking control method for the logistics area as described in any one of claims 1-6.
Citation Information
Patent Citations
Smoking behavior detection method and device, computer equipment and storage medium
CN111488841A
Smoking behavior identification method and device
CN112818919A
Detection method and device, detection equipment and storage medium
CN112950563A
Method and device for detecting smoking behavior
CN114140878A
Video monitoring method and device, computer equipment and storage medium
CN114173094A