Security visual management system and method for intelligent building

By combining the surveillance images acquired by wide-angle and high-precision cameras, and extracting and fusing feature vectors for classifier judgment, the blind spots and insufficient manual patrols of the existing security system are solved, and panoramic real-time monitoring and accurate identification of smart buildings are realized.

CN120339930AInactive Publication Date: 2025-07-18RUIYING HI-TECH TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410086843.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-22
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing building security system, the access control system has limited identification capabilities and cannot fully cover the safe area. Manual patrols are time-consuming and labor-intensive and prone to fatigue. There are blind spots and human negligence, and real-time monitoring and timely response cannot be achieved.

Method used

A wide-angle camera is used to obtain panoramic surveillance images and a high-precision camera is used to obtain close-up surveillance images. A multi-scale panoramic correlation feature vector and facial area state feature vector are generated through feature extraction and encoding. After the fusion, the classifier is used to determine whether there are suspicious people entering and leaving the building.

Benefits of technology

The detection and identification capabilities of suspicious people entering and leaving buildings have been improved, the security management level of smart buildings has been enhanced, and all-round real-time monitoring and accurate identification have been achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339930A_ABST
    Figure CN120339930A_ABST
Patent Text Reader

Abstract

The invention relates to the field of security and protection visual management, and particularly discloses a security and protection visual management system and method for an intelligent building, and the method comprises the steps: firstly obtaining a panoramic monitoring image collected by a wide-angle camera and a close-range monitoring image collected by a high-precision camera; respectively carrying out feature extraction and feature coding on the panoramic monitoring image and the close-range monitoring image to obtain a multi-scale panoramic association feature vector and a facial region state feature vector; and the multi-scale panoramic association feature vector and the facial region state feature vector are fused and a classification result is obtained through a classifier, and the classification result is used for judging whether a suspicious person enters and exits the building, so that the capability of detecting and identifying the suspicious person entering and exiting the building is improved. And the security management level of the intelligent building is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of security visual management, and more specifically, to a security visual management system and method for smart buildings. Background Art

[0002] Security refers to a series of measures and technical means taken to protect the safety of personnel, property, and information. It includes access control systems, surveillance systems, alarm systems, etc. By restricting access, monitoring in real time, and giving timely alarms, the security prevention ability can be improved, criminal acts can be prevented, and emergencies can be quickly responded to ensure the safety of personnel and property.

[0003] Existing buildings usually adopt a combination of access control system recognition and manual patrol to identify abnormal behaviors and intrusion events. However, the recognition ability of the access control system is limited and can only monitor specific areas or specific scenarios, unable to comprehensively cover the entire security area, resulting in the risk of blind spots. The manual patrol method is time-consuming and laborious and prone to fatigue, with the possibility of human negligence or missed inspections, unable to achieve real-time monitoring and timely response. In addition, relying solely on manual patrol is easily affected by subjective factors, and the recognition accuracy and consistency fluctuate.

[0004] Therefore, a security visual management system and method for smart buildings are desired. Summary of the Invention

[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide a security visual management system and method for smart buildings, which first obtain panoramic surveillance images collected by wide-angle cameras and close-up surveillance images collected by high-precision cameras, and then perform feature extraction and feature encoding on the panoramic surveillance images and the close-up surveillance images respectively to obtain multi-scale panoramic correlation feature vectors and facial region status feature vectors. Finally, the multi-scale panoramic correlation feature vectors and the facial region status feature vectors are fused and passed through a classifier to obtain a classification result, and the classification result is used to determine whether there are suspicious personnel entering or leaving the building, thereby improving the detection and recognition ability of suspicious personnel entering or leaving the building and enhancing the security management level of smart buildings.

[0006] According to one aspect of this application, a security visual management system for smart buildings is provided, which includes: A surveillance image acquisition module for acquiring panoramic surveillance images collected by wide-angle cameras and multiple close-up surveillance images collected by high-precision cameras; A surveillance image feature extraction module for extracting multi-scale panoramic correlation feature vectors and facial region status feature vectors from the panoramic surveillance images collected by wide-angle cameras and the multiple close-up surveillance images collected by high-precision cameras; An anti-theft feature fusion module, configured to fuse the multi-scale panoramic correlation feature vector and the facial area status feature vector to obtain an abnormal behavior classification feature matrix; A suspicious person judgment result generation unit, configured to obtain a classification result by passing the abnormal behavior classification feature matrix through a classifier, where the classification result is used to determine whether there are suspicious persons entering or leaving the building.

[0007] According to another aspect of the present application, there is provided a security visualization management method for a smart building, including: Obtaining a panoramic surveillance image collected by a wide-angle camera and a plurality of close-up surveillance images collected by a high-precision camera; Extracting a multi-scale panoramic correlation feature vector and a facial area status feature vector from the panoramic surveillance image collected by the wide-angle camera and the plurality of close-up surveillance images collected by the high-precision camera; Fusing the multi-scale panoramic correlation feature vector and the facial area status feature vector to obtain an abnormal behavior classification feature matrix; Passing the abnormal behavior classification feature matrix through a classifier to obtain a classification result, where the classification result is used to determine whether there are suspicious persons entering or leaving the building.

[0008] Compared with the prior art, a security visualization management system and method for a smart building provided by the present application first obtains a panoramic surveillance image collected by a wide-angle camera and a close-up surveillance image collected by a high-precision camera, and then performs feature extraction and feature encoding on the panoramic surveillance image and the close-up surveillance image respectively to obtain a multi-scale panoramic correlation feature vector and a facial area status feature vector. Finally, the multi-scale panoramic correlation feature vector and the facial area status feature vector are fused and passed through a classifier to obtain a classification result, where the classification result is used to determine whether there are suspicious persons entering or leaving the building, thereby improving the detection and recognition ability of suspicious persons entering or leaving the building and enhancing the security management level of the smart building. Description of the Drawings

[0009] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application, and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0010] Figure 1 It is a block diagram of a security visualization management system for a smart building according to an embodiment of the present application.

[0011] Figure 2It is a block diagram of a monitoring image feature extraction module in a security visual management system for smart buildings according to an embodiment of the present application.

[0012] Figure 3 It is a schematic diagram of the architecture of a security visual management system for smart buildings according to an embodiment of the present application.

[0013] Figure 4 It is a flowchart of a security visual management method for smart buildings according to an embodiment of the present application.

[0014] Figure 5 It is a block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0015] Next, exemplary embodiments according to the present application will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0016] Exemplary system Figure 1 It is a block diagram of a security visual management system for smart buildings according to an embodiment of the present application. As Figure 1 shown, a security visual management system 100 for smart buildings according to an embodiment of the present application includes: a monitoring image acquisition module 110, configured to acquire a panoramic monitoring image collected by a wide-angle camera and a plurality of close-up monitoring images collected by a high-precision camera; a monitoring image feature extraction module 120, configured to extract a multi-scale panoramic correlation feature vector and a facial area status feature vector from the panoramic monitoring image collected by the wide-angle camera and the plurality of close-up monitoring images collected by the high-precision camera; a security feature fusion module 130, configured to fuse the multi-scale panoramic correlation feature vector and the facial area status feature vector to obtain an abnormal behavior classification feature matrix; and a suspicious person judgment result generation unit 140, configured to pass the abnormal behavior classification feature matrix through a classifier to obtain a classification result, and the classification result is used to judge whether there are suspicious persons entering or leaving the building.

[0017] In the above-mentioned security visual management system 100 for intelligent buildings, the monitoring image acquisition module 110 is used to acquire panoramic monitoring images collected by wide-angle cameras and multiple close-up monitoring images collected by high-precision cameras. It should be understood that wide-angle cameras usually have a wider field of view, can cover a larger area, and provide an overall monitoring perspective. High-precision cameras, on the other hand, can provide higher resolution and clearer images, capable of capturing more details and monitoring information of specific areas. By combining wide-angle cameras and high-precision cameras, the dual advantages of panoramic images and close-up images can be obtained. Specifically, panoramic images can help monitoring personnel gain an overall understanding of the scene, understand the situation of the entire monitoring area, and quickly detect abnormal situations, while close-up images can provide clearer details, such as facial features, etc., which are helpful for more accurate identification and analysis. Combining panoramic monitoring images and close-up monitoring images can provide more comprehensive and detailed monitoring coverage. Using these two types of images comprehensively can enhance the security of the monitoring system and improve the perception and response capabilities to potential threats.

[0018] In the above-mentioned security visual management system 100 for intelligent buildings, the monitoring image feature extraction module 120 is used to extract multi-scale panoramic correlation feature vectors and facial region state feature vectors from the panoramic monitoring images collected by wide-angle cameras and the multiple close-up monitoring images collected by high-precision cameras. It should be understood that panoramic monitoring images are collected by wide-angle cameras and have a wider field of view, capable of capturing a larger range of the monitoring area. However, due to the wide field of view of wide-angle cameras, the target objects in the images may appear smaller. To improve the recognition and tracking capabilities of target objects, multi-scale feature vectors can be extracted from panoramic images. This can analyze the target at different scales, thus providing more comprehensive and accurate monitoring results. The close-up monitoring images collected by high-precision cameras have higher resolution and better image quality, capable of providing more detailed information. In the monitoring scenario, the facial region is an important target for recognition and analysis. By extracting the state feature vectors of the facial region from close-up images, such as expressions, eye gazes, postures, etc., it can be used for applications such as face recognition, emotion analysis, and behavior recognition. This can more accurately judge the state and behavior of personnel and provide more refined monitoring and security analysis. Extracting multi-scale panoramic correlation feature vectors and facial region state feature vectors from the panoramic monitoring images collected by wide-angle cameras and the multiple close-up monitoring images collected by high-precision cameras can improve the analysis and recognition capabilities of the monitoring system and enhance the perception and understanding of target objects and personnel states.

[0019] Figure 2 It is a block diagram of the monitoring image feature extraction module in the security visual management system for intelligent buildings according to the embodiment of the present application. As Figure 2As shown, in a specific embodiment of the present application, the monitoring image feature extraction module 120 includes: a panoramic image feature extraction unit 121, configured to extract features from the panoramic monitoring image collected by the wide-angle camera to obtain a panoramic multi-channel image; a panoramic image feature encoding unit 122, configured to perform feature encoding on the panoramic multi-channel image to obtain the multi-scale panoramic correlation feature vector; a close-range image feature extraction unit 123, configured to obtain multiple face region of interest feature maps from the multiple close-range monitoring images collected by the high-precision camera through a target detection network; and a face region of interest feature encoding unit 124, configured to perform feature encoding on the multiple face region of interest feature maps to obtain the face region state feature vector. It should be understood that a panoramic multi-channel image refers to dividing the panoramic image into multiple channels during the feature extraction process, and each channel represents different image features. Such a representation method can separate different feature information and provide a richer visual expression. For example, color channels, texture channels, edge channels, etc. can be used to represent different features. Through the panoramic multi-channel image, more information can be provided for object recognition, tracking, and analysis. Different channels can capture different image features, which can improve the recognition accuracy and robustness of the object, and at the same time provide more features for behavior analysis and event detection. To convert the preprocessed panoramic monitoring image into a multi-channel image, different methods can be used to generate the multi-channel image. One common method is to use a filter bank. The filter bank can contain filters with different directions and scales. By performing a convolution operation on the panoramic image, feature response maps of multiple channels can be obtained, and each channel corresponds to a specific image feature, such as edges, textures, colors, etc. Further, the panoramic monitoring image usually contains a large amount of details and multi-scale information. By performing multi-scale feature encoding on the panoramic multi-channel image, image features at different scales can be captured, thereby providing a more comprehensive and rich image representation to represent the image in a compact vector form for subsequent analysis and processing.

[0020] Furthermore, in order to extract the facial information in each surveillance image and represent it in the form of a specific feature map, and considering that when judging and recognizing multiple close-range surveillance images collected by the high-precision camera, the hidden features of the people entering and leaving should be focused on. Therefore, if the useless interference feature information can be filtered out when mining the features of the multiple close-range surveillance images collected by the high-precision camera, it is obvious that the accuracy of judging suspicious people can be improved. Based on this, in the technical solution of this application, the multiple close-range surveillance images collected by the high-precision camera are further passed through a target detection network to obtain multiple facial region of interest feature maps. Here, the target detection network is an anchor-window-based target detection network. Specifically, the target anchoring layer of the target detection network is used to slide with the anchor box B to process the multiple close-range surveillance images collected by the high-precision camera, so as to frame the facial region of interest in the close-range surveillance image, thereby obtaining the multiple facial region of interest feature maps. Correspondingly, in a specific example of this application, the anchor-window-based target detection network is Fast R-CNN, Faster R-CNN or RetinaNet. More specifically, the following target detection formula is used to process the multiple close-range surveillance images collected by the high-precision camera to obtain the multiple facial region of interest feature maps; Among them, the target detection formula is: Among them, is the multiple close-range surveillance images collected by the high-precision camera, is the anchor box, is the facial region of interest feature map, represents classification, represents regression.

[0021] In particular, feature encoding can extract the key information in the facial region of interest through learning, so that the feature maps of different facial regions of interest can be converted into feature vectors with similar representation forms. This enables comparison and matching between different faces, thus realizing tasks such as face recognition.

[0022] In a specific embodiment of the present application, the panoramic image feature extraction unit 121 includes: extracting a Local Binary Patterns (LBP) map and a Canny edge detection feature map from the panoramic surveillance image collected by the wide-angle camera; combining the panoramic surveillance image collected by the wide-angle camera, the LBP map, and the Canny edge detection feature map to obtain the panoramic multi-channel image. It should be understood that LBP (Local Binary Patterns) is a method for describing local texture features of an image. By calculating the gray-scale difference between the neighboring pixels around each pixel and the central pixel and encoding the difference as a binary pattern, an LBP image can be obtained. The LBP image can effectively represent the texture features of the image and is very useful for identifying and distinguishing different texture patterns. Further, Canny edge detection is a classic edge detection algorithm that can extract edge information in an image. By calculating the gradient magnitude and direction of the pixels in the image and applying non-maximum suppression and double-threshold processing, the Canny algorithm can achieve accurate detection of edges. Extracting the Canny edge detection feature map can capture the boundary and contour information in the image. By extracting the LBP map and the Canny edge detection feature map from the panoramic surveillance image, the local texture and edge information of the image can be extracted to help identify and analyze different target objects in the image. Furthermore, the panoramic surveillance image provides visual information of the overall scene and can be used for positioning and overall analysis. The LBP map can provide local texture information of the image and is very useful for texture analysis and target recognition, while the Canny edge detection feature map can provide edge information of the image and helps in target contour extraction and edge detection. By combining the LBP map and the Canny edge detection feature map with the panoramic surveillance image, different visual information can be captured at multiple scales, which helps to improve the robustness and accuracy of target detection and analysis. Among them, if the sizes of the panoramic surveillance image collected by the wide-angle camera, the LBP map, and the Canny edge detection feature map do not match, they can be adjusted or cropped to obtain the same width and height. Then, normalization processing is performed on the LBP map and the Canny edge detection feature map to map the pixel value range of the LBP map and the Canny edge detection feature map to between 0 and 255. This can be achieved by linear scaling or using a suitable normalization method. Then, a new multi-channel image is created, with the panoramic surveillance image as one channel and the LBP map and the Canny edge detection feature map as other channels. More specifically, the panoramic surveillance image collected by the wide-angle camera, the LBP map, and the Canny edge detection feature map are combined according to the following cascading formula to obtain the panoramic multi-channel image; Among them, the cascade formula is: Among them, represents the panoramic monitoring image collected by the wide-angle camera, represents the LBP local binary pattern map, represents the Canny edge detection feature map, represents the panoramic multi-channel image, represents the cascade function.

[0023] In a specific embodiment of the present application, the panoramic image feature encoding unit 122 includes: passing the panoramic multi-channel image through a pyramid network to obtain a panoramic feature pyramid; performing upsampling on the panoramic feature pyramid to obtain a panoramic detection feature map; passing the panoramic detection feature map through dilated convolutional kernels with different dilation rates to obtain the multi-scale panoramic correlation feature vector. It should be understood that the pyramid network can generate feature images of different scales and is a multi-level network structure. Each layer performs different degrees of downsampling on the input image to capture features at different scales. At the higher levels of the pyramid, the feature images have lower resolutions but can capture broader scene information; while at the lower levels, the feature images have higher resolutions and can provide more detailed information. Through the panoramic feature pyramid, the panoramic image can be analyzed at different scales, thereby improving the robustness and accuracy of object detection. Specifically, a shallow feature matrix is extracted from the i-th layer of the pyramid network, where the i-th layer is the first to the sixth layer of the pyramid network; a deep feature matrix is extracted from the j-th layer of the pyramid network, and the ratio between the j-th layer and the i-th layer is greater than or equal to 5; a shallow and deep feature fusion module is used to fuse the shallow feature matrix and the deep feature matrix to obtain the panoramic feature pyramid.

[0024] Furthermore, by performing upsampling on the panoramic feature pyramid, the resolution of the feature image can be restored. Specifically, upsampling can be achieved through interpolation methods, such as bilinear interpolation or deconvolution operations. More specifically, the low-resolution feature map is enlarged to the same size as the original panoramic image. The panoramic detection feature map with restored resolution can better capture the details and edge information of the target, improving the accuracy of object detection.

[0025] Furthermore, by introducing dilated sampling on the input feature map, the dilated convolutional kernel can perceive the panoramic detection feature map at different scales. A smaller dilation rate can capture detailed information, while a larger dilation rate can cover a wider area to improve the detection and recognition capabilities of targets at different scales. In addition, since a smaller dilation rate can retain more detailed information but has a larger computational cost, while a larger dilation rate can reduce the computational cost but may lose some details, by selecting appropriate dilation rates at different scales, the computational complexity can be reduced while maintaining a certain level of accuracy. In summary, by using dilated convolutional kernels with different dilation rates, multi-scale panoramic correlation feature vectors can be obtained, thereby improving the perception and expression capabilities of targets in panoramic images. Specifically, each layer using a dilated convolutional kernel with a first dilation rate performs dilated convolution processing based on the dilated convolutional kernel with the first dilation rate, pooling processing along the channel dimension, and non-linear activation processing on the input data during the forward pass of the layer to output a first feature matrix from the last layer of the dilated convolutional model with the first dilation rate; each layer using a dilated convolutional kernel with a second dilation rate performs dilated convolution processing based on the dilated convolutional kernel with the second dilation rate, pooling processing along the channel dimension, and non-linear activation processing on the input data during the forward pass of the layer to output a second feature matrix from the last layer of the dilated convolutional model with the second dilation rate; the first feature matrix and the second feature matrix are fused and pooled to obtain a multi-scale panoramic correlation feature vector.

[0026] In a specific embodiment of the present application, the facial region of interest feature encoding unit 124 includes: arranging the multiple facial region of interest feature maps to obtain a three-dimensional input tensor of the facial region; passing the three-dimensional input tensor of the facial region through a facial region state feature extractor with a three-dimensional convolution kernel to obtain a facial region state feature map; pooling the facial region state feature map to obtain the facial region state feature vector. It should be understood that the facial region of interest feature maps are usually obtained by segmenting a facial image or detecting key points. Each feature map corresponds to a specific facial region, such as eyes, nose, mouth, etc. Arranging these feature maps can integrate the information of each facial region into a unified data structure for subsequent processing. Arranging multiple facial region of interest feature maps can form a three-dimensional input tensor, where the dimensions can represent different facial regions, feature channels, and spatial positions. This three-dimensional representation can better capture the relationships and structures between facial regions, helping to extract richer facial features. Further, the facial region state feature map is obtained by performing a convolution operation on the three-dimensional input tensor of the facial region. Through the convolution operation, the information of surrounding pixels can be integrated into each pixel point, thereby obtaining a richer feature representation. This helps to improve the representational ability of the facial region and capture more detailed and structural information of the facial region. Specifically, each layer of the facial region state feature extractor with a three-dimensional convolution kernel respectively performs the following operations on the input data during the forward pass of the layer: using the convolution units of each layer to perform convolution processing on the input data based on the three-dimensional convolution kernel to obtain a convolution feature map; passing the convolution feature map through the spatial attention units of each layer to obtain a spatial attention map; calculating the element-wise multiplication of the convolution feature map and the spatial attention map to obtain a spatial attention feature map; inputting the spatial attention feature map into the non-linear activation units of each layer to obtain an activation feature map; wherein, the output of the last layer of the facial region state feature extractor with a three-dimensional convolution kernel is the facial region state feature map.

[0027] Furthermore, feature maps usually have a high dimension, where each position contains a large amount of feature information. However, in some cases, more attention is paid to the overall features of the entire facial region rather than the details of each position. By performing a pooling operation on the facial region state feature map, it can be reduced to a facial region state feature vector, where each element represents a certain overall feature of the entire facial region.

[0028] In the above-mentioned security visualization management system 100 for smart buildings, the security feature fusion module 130 is used to fuse the multi-scale panoramic correlation feature vector and the facial area status feature vector to obtain an abnormal behavior classification feature matrix. It should be understood that the multi-scale panoramic correlation feature vector and the facial area status feature vector capture facial information at different levels and angles respectively. Fusing these two features can comprehensively utilize the information they contain and provide a more comprehensive and rich description of facial features.

[0029] In the above-mentioned security visualization management system 100 for smart buildings, the suspicious person judgment result generation unit 140 is used to pass the abnormal behavior classification feature matrix through a classifier to obtain a classification result, and the classification result is used to judge whether there are suspicious persons entering or leaving the building. It should be understood that the classifier can learn according to the features and labels of the training data, so as to be able to provide relatively accurate and consistent judgment results. Compared with manual judgment, the classifier can better capture the patterns and features of abnormal behaviors and improve the accuracy of judgment. Specifically, it is judged whether there are suspicious persons entering or leaving the building according to the classification result. For example, if the classification result is "abnormal", it can be considered that there are suspicious persons.

[0030] In a specific embodiment of the present application, there is also a training module for training the pyramid network, the dilated convolutional kernels with different dilation rates, the object detection network, the facial area status feature extractor using a three-dimensional convolutional kernel, and the classifier.

[0031] In a specific embodiment of the present application, the training module includes: a training data acquisition unit for acquiring training data, where the training data includes training panoramic surveillance images collected by a wide-angle camera and training close-up surveillance images collected by a high-precision camera; a training data extraction unit for extracting a training LBP (Local Binary Pattern) map and a training Canny edge detection feature map of the training panoramic surveillance images collected by the wide-angle camera; a training data merging unit for merging the training panoramic surveillance images collected by the wide-angle camera, the training LBP map, and the training Canny edge detection feature map to obtain a training panoramic multi-channel image; a training pyramid unit for passing the training panoramic multi-channel image through a pyramid network to obtain a training panoramic feature pyramid; a panoramic feature upsampling unit for upsampling the training panoramic feature pyramid to obtain a training panoramic detection feature map; a training multi-scale unit for passing the training panoramic detection feature map through dilated convolutional kernels with different dilation rates to obtain a training multi-scale panoramic correlation feature vector; a training data object detection unit for passing the training close-up surveillance images collected by the high-precision camera through an object detection network to obtain a training facial region of interest feature map; an arrangement unit for arranging the training facial region of interest feature map to obtain a training facial region three-dimensional input tensor; a training facial feature extraction unit for passing the training facial region three-dimensional input tensor through a facial region state feature extractor using a three-dimensional convolutional kernel to obtain a training facial region state feature map; a training facial feature pooling unit for pooling the training facial region state feature map to obtain the training facial region state feature vector; a training fusion unit for fusing the training multi-scale panoramic correlation feature vector and the training facial region state feature vector to obtain an abnormal behavior classification feature matrix; a classification loss unit for passing the abnormal behavior classification feature matrix through the classifier to obtain a classification loss function value; a calculation unit for calculating the inter-node topological dimension information measure factor between the training multi-scale panoramic correlation feature vector and the training facial region state feature vector; a loss function unit for using the weighted sum of the inter-node topological dimension information measure factor and the classification loss function value as the loss function value to train the pyramid network, the dilated convolutional kernels with different dilation rates, the object detection network, the facial region state feature extractor using a three-dimensional convolutional kernel, and the classifier.

[0032] In particular, in the technical solution of this application, when fusing and training the multi-scale panoramic correlation feature vector and the training facial region state feature vector, there may be a situation where the feature scales between the high-dimensional implicit feature distribution of the training multi-scale panoramic correlation feature vector and the high-dimensional implicit correlation feature of the training facial region state feature vector are different, and the feature distribution manifolds corresponding to them in the high-dimensional feature space are also different. Specifically, the training multi-scale panoramic correlation feature vector is usually obtained by processing and analyzing panoramic surveillance images, and it may contain features related to panoramic scene perception, target tracking, etc. For example, the panoramic scene perception features may include object or texture information at multiple scales, and this information presents diversity and complexity in the feature space. The training facial region state feature vector is obtained by processing the facial region in the close-range surveillance image, and it may contain features related to face recognition, expression analysis, etc. For example, the facial region state features may include face key points, expression features, etc., and these features present a relatively regular distribution in the feature space. Since the training multi-scale panoramic correlation feature vector and the training facial region state feature vector have different feature scales and feature distribution manifolds, simply weighting and summing them may not effectively fuse their feature information. Based on this, in the technical solution of this application, when fusing the training multi-scale panoramic correlation feature vector and the training facial region state feature vector, the inter-node topological dimension information metric factor between the training multi-scale panoramic correlation feature vector and the training facial region state feature vector is calculated as the loss function value to solve the problem that the feature scales between the high-dimensional implicit feature distribution of the training multi-scale panoramic correlation feature vector and the high-dimensional implicit correlation feature of the training facial region state feature vector are different, and the feature distribution manifolds corresponding to them in the high-dimensional feature space are also different, so the feature information of the two cannot be fused by the method of weighting and summing.

[0033] Specifically, the inter-node topological dimension information metric factor between the training multi-scale panoramic correlation feature vector and the training facial region state feature vector is calculated according to the following metric factor formula; wherein, the metric factor formula is: where represents the training multi-scale panoramic correlation feature vector, represents the training facial region state feature vector, and are column vectors, represents the transposed vector of the training multi-scale panoramic correlation feature vector, represents calculating the vector and the vector covariance, represents calculating the variance of the vector, Represents the information metric factor of the topological dimension between nodes.

[0034] More specifically, the covariance between vectors is calculated using the following covariance formula; where the covariance formula is: Where the covariance formula is: Where and are the eigenvalues at the th position of the training multi-scale panoramic correlation feature vector and the training facial region state feature vector respectively, and and are the mean values of the eigenvalues at each position of the training multi-scale panoramic correlation feature vector and the training facial region state feature vector respectively, represents the length of the feature vector, represents calculating the covariance of the vectors and the vector .

[0035] Furthermore, the variance of the vector is calculated using the following variance formula; Where the variance formula is: Where the variance formula is: Where represents the feature vector, represents the th position eigenvalue of the feature vector , represents the mean value of the eigenvalues at each position of the feature vector , represents the length of the feature vector, calculates the variance of the vector .

[0036] That is, when using the trained abnormal behavior classification feature matrix of the trained multi-scale panoramic correlation feature vector and the trained facial region state feature vector as a matrix for classification, in order to avoid the increase in the sparsity of the trained abnormal behavior classification feature matrix caused by the feature incompatibility between the trained multi-scale panoramic correlation feature vector and the trained facial region state feature vector, that is, there is a certain conflict or contradiction between the two feature vectors, and this information is harmful to the classification expression of the trained abnormal behavior classification feature matrix, but it is not explicitly reflected in the original feature vectors. Therefore, in order to eliminate or reduce this feature incompatibility, in the technical solution of this application, a method based on the inter-node topological dimension information metric factor is further introduced. This method can calculate the inter-node topological dimension information metric factor between the trained multi-scale panoramic correlation feature vector and the trained facial region state feature vector, and use the inter-node topological dimension information metric factor as a loss function for training. Specifically, by using the inter-node topological dimension information metric factor, that is, each element in the feature vector can express its topological structure or topological relationship in the high-dimensional space, so as to perform the topological dimension information of the feature vector, so as to realize the correlation migration between sub-regions and within sub-regions of the encoded feature vector in the feature space, that is, the feature vectors between different sub-regions can influence and adjust each other, so as to promote the bidirectional approximation between feature vectors, that is, the correlation expression between feature vectors is closer and more consistent, so as to meet the density of the feature expression of the trained abnormal behavior classification feature matrix under the probability density, that is, each element in the trained abnormal behavior classification feature matrix can reflect the correlation and dependence between the two feature vectors, so as to improve the classification accuracy.

[0037] Figure 3 Schematic diagram of the architecture of the security visualization management system for smart buildings according to an embodiment of the present application. As Figure 3As shown, the security visual management system 100 for smart buildings according to an embodiment of the present application includes: First, obtain the panoramic surveillance images collected by the wide-angle camera and the close-up surveillance images collected by the high-precision camera. Then, extract the LBP local binary pattern map and the Canny edge detection feature map of the panoramic surveillance images collected by the wide-angle camera. Next, merge the panoramic surveillance images collected by the wide-angle camera, the LBP local binary pattern map, and the Canny edge detection feature map to obtain a panoramic multi-channel image. Immediately afterwards, pass the panoramic multi-channel image through a pyramid network to obtain a panoramic feature pyramid. Then, upsample the panoramic feature pyramid to obtain a panoramic detection feature map. Further, pass the panoramic detection feature map through dilated convolutional kernels with different dilation rates to obtain multi-scale panoramic correlation feature vectors. Furthermore, pass the close-up surveillance images collected by the high-precision camera through an object detection network to obtain a face region of interest feature map. Subsequently, arrange the face region of interest feature map to obtain a three-dimensional input tensor of the face region. Then, pass the three-dimensional input tensor of the face region through a face region state feature extractor using a three-dimensional convolutional kernel to obtain a face region state feature map. Next, pool the face region state feature map to obtain a face region state feature vector. In addition, fuse the multi-scale panoramic correlation feature vectors and the face region state feature vectors to obtain an abnormal behavior classification feature matrix. Finally, pass the abnormal behavior classification feature matrix through a classifier to obtain a classification result, and the classification result is used to determine whether there are suspicious persons entering or leaving the building.

[0038] In summary, the embodiment of the present application first obtains the panoramic surveillance images collected by the wide-angle camera and the close-up surveillance images collected by the high-precision camera, then uses deep learning technology to perform feature extraction and correlation analysis on the two, and finally obtains a classification result through a classifier to determine whether there are suspicious persons entering or leaving the building.

[0039] As described above, the security visual management system 100 for smart buildings according to an embodiment of the present application can be implemented in various terminal devices, such as a server deployed with a security visual management control algorithm for smart buildings. In one example, the security visual management system 100 for smart buildings can be integrated into a terminal device as a software module and / or a hardware module. For example, the security visual management system 100 for smart buildings can be a software module in the operating system of the terminal device, or can be an application program developed for the terminal device; of course, the security visual management system 100 for smart buildings can also be one of the many hardware modules of the terminal device.

[0040] Alternatively, in another example, the security visual management system 100 for smart buildings and the terminal device can also be separate devices, and the security visual management system 100 for smart buildings can be connected to the terminal device through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.

[0041] Exemplary method Figure 4 FIG. is a flowchart of a security visual management method for smart buildings according to an embodiment of the present application. As Figure 4 shown, the security visual management method for smart buildings according to an embodiment of the present application includes: S110, obtaining a panoramic surveillance image collected by a wide-angle camera and a plurality of close-up surveillance images collected by a high-precision camera; S120, extracting a multi-scale panoramic correlation feature vector and a facial region status feature vector from the panoramic surveillance image collected by the wide-angle camera and the plurality of close-up surveillance images collected by the high-precision camera; S130, fusing the multi-scale panoramic correlation feature vector and the facial region status feature vector to obtain an abnormal behavior classification feature matrix; S140, passing the abnormal behavior classification feature matrix through a classifier to obtain a classification result, where the classification result is used to determine whether there are suspicious persons entering or leaving the building.

[0042] Here, those skilled in the art can understand that the specific operations of each step in the above security visual management method for smart buildings have been described in detail in the description of the security visual management system for smart buildings above with reference to Figures 1 to 3 and thus, the repeated description thereof will be omitted.

[0043] Exemplary electronic device Based on the above security visual management method for smart buildings, the present invention also proposes an electronic device.

[0044] Figure 5 FIG. is a structural block diagram of an electronic device according to an embodiment of the present invention.

[0045] As Figure 5 shown, the electronic device 10 includes: a processor 11 and a memory 13. Among them, the processor 11 and the memory 13 are connected, such as through a bus 12. Optionally, the electronic device 10 may further include a transceiver 14. It should be noted that in practical applications, the transceiver 14 is not limited to one, and the structure of the electronic device 10 does not constitute a limitation on the embodiments of the present invention.

[0046] The processor 11 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present invention. The processor 11 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0047] The bus 12 may include a path for transmitting information between the above components. The bus 12 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 12 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0048] The memory 13 is used to store a computer program corresponding to the security visual management method for intelligent buildings in the above embodiments of the present invention, and this computer program is controlled and executed by the processor 11. The processor 11 is used to execute the computer program stored in the memory 13 to implement the content shown in the foregoing method embodiments.

[0049] Among them, the electronic device 10 includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The illustrated electronic device 10 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0050] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus or device), or in combination with these instruction execution systems, apparatuses or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by or in combination with an instruction execution system, apparatus or device. More specific examples (non-exhaustive list) of the computer-readable medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting or otherwise processing it as appropriate, and then storing it in a computer memory.

[0051] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0052] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0053] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0054] In the present invention, unless otherwise clearly stipulated and defined, terms such as "installed", "connected", "joined", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention may be understood according to specific circumstances.

[0055] In the present invention, unless otherwise clearly stipulated and defined, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may mean that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "underneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.

[0056] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A security visualization management system for smart buildings, characterized in that, Including: A monitoring image acquisition module, configured to acquire a panoramic monitoring image collected by a wide-angle camera and a plurality of close-range monitoring images collected by a high-precision camera; A monitoring image feature extraction module, configured to extract a multi-scale panoramic correlation feature vector and a facial area status feature vector from the panoramic monitoring image collected by the wide-angle camera and the plurality of close-range monitoring images collected by the high-precision camera; A security feature fusion module, configured to fuse the multi-scale panoramic correlation feature vector and the facial area status feature vector to obtain an abnormal behavior classification feature matrix; A suspicious person judgment result generation unit, configured to pass the abnormal behavior classification feature matrix through a classifier to obtain a classification result, and the classification result is used to judge whether there are suspicious persons entering or leaving the building.

2. The security visual management system for intelligent buildings according to claim 1, wherein The monitoring image feature extraction module includes: A panoramic image feature extraction unit, configured to perform feature extraction on the panoramic monitoring image collected by the wide-angle camera to obtain a panoramic multi-channel image; A panoramic image feature encoding unit, configured to perform feature encoding on the panoramic multi-channel image to obtain the multi-scale panoramic correlation feature vector; A close-range image feature extraction unit, configured to pass the plurality of close-range monitoring images collected by the high-precision camera through an object detection network to obtain a plurality of facial region of interest feature maps; A facial region of interest feature encoding unit, configured to perform feature encoding on the plurality of facial region of interest feature maps to obtain the facial area status feature vector.

3. The security visual management system for smart buildings according to claim 2, wherein The panoramic image feature extraction unit includes: Extract a local binary pattern (LBP) map and a Canny edge detection feature map from the panoramic monitoring image collected by the wide-angle camera; Merge the panoramic monitoring image collected by the wide-angle camera, the LBP map, and the Canny edge detection feature map to obtain the panoramic multi-channel image.

4. The security visual management system for smart buildings according to claim 3, wherein The panoramic image feature encoding unit includes: Pass the panoramic multi-channel image through a pyramid network to obtain a panoramic feature pyramid; Perform upsampling on the panoramic feature pyramid to obtain a panoramic detection feature map; Pass the panoramic detection feature map through dilated convolutional kernels with different dilation rates to obtain the multi-scale panoramic correlation feature vector.

5. The security visual management system for intelligent buildings according to claim 4, wherein, The facial region of interest feature encoding unit includes: Arrange the plurality of facial region of interest feature maps to obtain a three-dimensional input tensor for the facial area; Pass the three-dimensional input tensor for the facial area through a facial area status feature extractor with a three-dimensional convolutional kernel to obtain a facial area status feature map; Perform pooling on the facial area status feature map to obtain the facial area status feature vector.

6. The security visual management system for smart buildings according to claim 5, wherein It further includes a training module for training the pyramid network, the dilated convolutional kernels with different dilation rates, the object detection network, the facial area status feature extractor with a three-dimensional convolutional kernel, and the classifier.

7. The security visual management system for intelligent buildings according to claim 6, characterized in that, The training module includes: A training data acquisition unit, configured to acquire training data, where the training data includes a training panoramic monitoring image collected by a wide-angle camera and a training close-range monitoring image collected by a high-precision camera; A training data extraction unit, configured to extract a training LBP (Local Binary Pattern) map and a training Canny edge detection feature map of the training panoramic surveillance image collected by the wide-angle camera; A training data merging unit, configured to merge the training panoramic surveillance image collected by the wide-angle camera, the training LBP map, and the training Canny edge detection feature map to obtain a training panoramic multi-channel image; A training pyramid unit, configured to obtain a training panoramic feature pyramid by passing the training panoramic multi-channel image through a pyramid network; A panoramic feature upsampling unit, configured to upsample the training panoramic feature pyramid to obtain a training panoramic detection feature map; A training multi-scale unit, configured to obtain a training multi-scale panoramic correlation feature vector by passing the training panoramic detection feature map through dilated convolution kernels with different dilation rates; A training data object detection unit, configured to obtain a training face region of interest feature map by passing the training close-range surveillance image collected by the high-precision camera through an object detection network; An arrangement unit, configured to arrange the training face region of interest feature map to obtain a training face region three-dimensional input tensor; A training face feature extraction unit, configured to obtain a training face region state feature map by passing the training face region three-dimensional input tensor through a face region state feature extractor with a three-dimensional convolution kernel; A training face feature pooling unit, configured to pool the training face region state feature map to obtain the training face region state feature vector; A training fusion unit, configured to fuse the training multi-scale panoramic correlation feature vector and the training face region state feature vector to obtain an abnormal behavior classification feature matrix; A classification loss unit, configured to obtain a classification loss function value by passing the abnormal behavior classification feature matrix through the classifier; A calculation unit, configured to calculate an inter-node topological dimension information metric factor between the training multi-scale panoramic correlation feature vector and the training face region state feature vector; A loss function unit, configured to use a weighted sum of the inter-node topological dimension information metric factor and the classification loss function value as a loss function value to train the pyramid network, the dilated convolution kernels with different dilation rates, the object detection network, the face region state feature extractor with a three-dimensional convolution kernel, and the classifier.

8. The security visual management system for intelligent buildings according to claim 7, wherein The calculation unit includes: Calculating the inter-node topological dimension information metric factor between the training multi-scale panoramic correlation feature vector and the training face region state feature vector according to the following metric factor formula; Among them, the metric factor formula is: Among them, represents the training multi-scale panoramic correlation feature vector, represents the training facial region state feature vector, and are column vectors, represents the transposed vector of the training multi-scale panoramic correlation feature vector, represents the calculation vector and the vector covariance, represents the variance of the calculation vector, represents the information metric factor of the topological dimension between nodes.

9. The security visual management system for intelligent buildings according to claim 8, characterized in that The calculation unit includes: Calculating the covariance between vectors according to the following covariance formula; Among them, the covariance formula is as follows: Among them, and are the eigenvalues at the -th position of the training multi-scale panoramic correlation feature vector and the training facial region state feature vector respectively, and and are the mean values of the eigenvalues at each position of the training multi-scale panoramic correlation feature vector and the training facial region state feature vector respectively, represents the length of the feature vector, represents calculating the covariance of the vector and the vector ; Calculating the variance of a vector according to the following variance formula; Among them, the variance formula is: Among them, represents the eigenvector, represents the eigenvector at the -th position of the eigenvalue, represents the average value of the eigenvalues at each position of the eigenvector , represents the length of the eigenvector, Calculate the vector variance.

10. A security visualization management method for smart buildings, characterized in that, Includes: Obtaining a panoramic surveillance image collected by a wide-angle camera and a plurality of close-range surveillance images collected by a high-precision camera; Extracting a multi-scale panoramic correlation feature vector and a face region state feature vector from the panoramic surveillance image collected by the wide-angle camera and the plurality of close-range surveillance images collected by the high-precision camera; Fuse the multi-scale panoramic correlation feature vector and the facial region state feature vector to obtain an abnormal behavior classification feature matrix; Pass the abnormal behavior classification feature matrix through a classifier to obtain a classification result, and the classification result is used to determine whether there are suspicious persons entering or leaving the building.