Security monitoring device supporting AI behavior analysis

Through multi-view camera groups and intelligent algorithms, efficient abnormal behavior detection and precise positioning in dynamic scenarios are achieved, and the shortcomings of existing security monitoring devices in dynamic scenario adaptability, multi-objective behavior correlation analysis and data security sharing are solved, and the intelligence and reliability of security monitoring devices are improved.

CN120279595APending Publication Date: 2025-07-08MEILIHUA INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510367599.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing security monitoring devices have shortcomings in dynamic scenario adaptability, multi-objective behavior correlation analysis capabilities, and data security and sharing mechanisms, which affects their practicality and reliability in complex scenarios.

Method used

A multi-view camera group is used to collect real-time video streams through the first camera and perform dynamic scene segmentation. The second camera is used to perform high-resolution supplementary acquisition. Combined with space-time alignment and fusion and frequency domain transformation, behavior feature vectors are extracted, and blockchain technology is used to achieve distributed update of reference behavior models to generate abnormal behavior heat maps.

Benefits of technology

It improves the accuracy and detection efficiency of behavior analysis in dynamic scenarios, enhances the ability of multi-objective behavior correlation analysis, and ensures data security and sharing through blockchain technology, improving the intelligence level and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention relates to the technical field of security and protection monitoring, in particular to a security and protection monitoring device supporting AI behavior analysis, and the device comprises the steps: collecting a real-time video stream of a target region through a first camera, carrying out the dynamic scene segmentation of the real-time video stream, and obtaining a plurality of sub-scene regions; if it is detected that the behavior mode of a certain sub-scene area deviates from a preset behavior benchmark, activating a second camera to perform high-resolution complementary collection; performing space-time alignment fusion on the video streams collected by the first camera and the second camera to generate an enhanced video stream; performing frequency domain transformation on each frame of image in the enhanced video stream, extracting a behavior feature vector, and calculating a similarity difference value between the behavior feature vector and the reference behavior model; and counting the maximum value and the minimum value of the similarity difference values of all the sub-scene areas in the same time period, and determining an abnormal behavior area in the target area based on the maximum value and the minimum value. The detection efficiency and the system intelligence level can be improved, and the system reliability is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of security monitoring, and specifically relates to a security monitoring device supporting AI behavior analysis. Background Art

[0002] With the rapid development of artificial intelligence technology, the application of AI behavior analysis in the field of security monitoring has become increasingly widespread. By combining AI algorithms and video monitoring technology, it is possible to achieve behavior recognition and anomaly warning of targets such as people and vehicles, thereby significantly improving the intelligence level and response efficiency of the security system. However, existing security monitoring devices still have certain deficiencies in supporting AI behavior analysis, which affects their practicality and reliability in complex scenarios.

[0003] After retrieval, the patent with the publication number CN115457449B provides an early warning system based on AI video analysis and security monitoring, and the publication date is March 24, 2023. The system includes an early warning object acquisition module, a real-time image acquisition module, an AI video analysis module, an early warning module, and a tracking module. By constructing a description information library and an identification reference information library, and using the AI video analysis method to analyze real-time monitoring images, the anomaly early warning function is realized, and it helps to improve the effectiveness of security monitoring by humans. However, this technical solution mainly relies on static description information libraries and identification reference information libraries. When facing dynamically changing scenarios, it may not be able to update model parameters in a timely manner to adapt to new behavior patterns, resulting in limitations in the accuracy and real-time performance of behavior analysis. In addition, this system does not fully consider the fusion analysis ability of multi-source data and is difficult to handle multi-target behavior interaction scenarios in complex environments.

[0004] After retrieval, the patent with the publication number CN111814646B proposes a monitoring method based on AI vision, and the publication date is April 5, 2024. This method performs pose detection on target users through a pose detection algorithm and identifies the staying position and duration of the target users, thereby realizing real-time monitoring and security warning of users. This technical solution can improve the detection efficiency and achieve high real-time performance, and is applicable to the field of intelligent security. However, this solution mainly focuses on single-target pose detection and behavior analysis, lacking the ability to deeply mine the relevance of multi-target behaviors. At the same time, the feedback mechanism for behavior analysis results of this system is relatively single, and it fails to make full use of emerging technologies such as blockchain to achieve secure storage and sharing of data, which may affect the overall reliability and scalability of the system.

[0005] The above problems indicate that there are still certain deficiencies in the existing security monitoring devices supporting AI behavior analysis in aspects such as dynamic scene adaptability, multi-object behavior correlation analysis ability, and data security and sharing mechanisms. Therefore, the present invention provides a security monitoring device supporting AI behavior analysis, aiming to optimize the behavior analysis ability in dynamic scenes, enhance the accuracy of multi-object behavior correlation analysis, and introduce data security and sharing mechanisms to meet the requirements of modern security fields for efficient and intelligent monitoring devices. Summary of the Invention

[0006] An embodiment of the present application provides a security monitoring device supporting AI behavior analysis, belonging to the technical field of security monitoring. The specific technical solution is as follows: To solve the above technical problems, the present application is proposed. An embodiment of the present application provides a security monitoring device supporting AI behavior analysis.

[0007] According to one aspect of the present application, there is provided a security monitoring device supporting AI behavior analysis, including: collecting a real-time video stream of a target area through a first camera; performing dynamic scene segmentation on the real-time video stream to obtain multiple sub-scene areas; wherein, the dynamic scene segmentation is based on the trajectory distribution of moving objects and the background change frequency in the target area; if it is detected that the behavior pattern of a certain sub-scene area deviates from a preset behavior benchmark, activating a second camera to perform high-resolution supplementary collection on the sub-scene area; wherein, the first camera and the second camera are respectively set at different viewing angle positions of the target area, and the field of view angle of the second camera covers the blind area of the first camera; performing spatio-temporal alignment and fusion on the video streams collected by the first camera and the second camera to generate an enhanced video stream; wherein, the spatio-temporal alignment and fusion adopts an adaptive weight distribution strategy, and the weight is determined by the clarity of the data collected by each camera and the trajectory consistency of moving objects; performing frequency domain transformation on each frame image in the enhanced video stream to extract behavior feature vectors; wherein, the frequency domain transformation formula is: Wherein, F(u, v) is the value at the frequency domain point (u, v), f(x, y) is the grayscale value of the pixel point (x, y), W is the image width, H is the image height, and i is the imaginary unit; calculate the similarity difference between the behavioral feature vector and the reference behavioral model; wherein, the reference behavioral model is trained and generated based on historical data and is updated distributively through blockchain technology; count the maximum and minimum values of the similarity differences of all sub-scene regions within the same time period; determine the abnormal behavior region in the target region based on the maximum and minimum values. In some embodiments, the background light source of the target region adopts a dimmable LED array, and the dimmable LED array dynamically adjusts the brightness according to the environmental light intensity; wherein, performing dynamic scene segmentation on the real-time video stream includes: calculating the optical flow vector of each pixel point in the real-time video stream; if the change in the direction and magnitude of the optical flow vector in a certain region exceeds a preset threshold, then mark this region as a dynamic sub-scene region.

[0008] In some embodiments, performing dynamic scene segmentation on the real-time video stream includes: performing edge detection on the real-time video stream to obtain an edge image; performing connected component analysis on the edge image to obtain a plurality of connected regions; if the area and shape features of a certain connected region meet the preset conditions, then mark it as a dynamic sub-scene region.

[0009] In some embodiments, performing spatio-temporal alignment and fusion on the video streams collected by the first camera and the second camera includes: performing geometric correction on the video streams of the first camera and the second camera based on the fixed markers in the target region; performing weighted fusion on the dynamic sub-scene regions in the video stream of the first camera and the corresponding regions in the video stream of the second camera to obtain a preliminary fusion region; performing denoising processing on the preliminary fusion region to obtain a final fusion region; generating the enhanced video stream according to the final fusion region.

[0010] In some embodiments, determining the abnormal behavior region in the target region based on the maximum and minimum values includes: if the ratio of the maximum and minimum values of the similarity differences within a certain time period exceeds a preset ratio, then mark the corresponding sub-scene region as an abnormal behavior region; generating a heat map of abnormal behaviors in the target region according to the spatial distribution of multiple abnormal behavior regions.

[0011] In some embodiments, the real-time video stream includes the daytime scene video stream and the nighttime scene video stream of the target region; performing dynamic scene segmentation on the real-time video stream includes: comparing the lighting distribution differences in the target region between the daytime scene video stream and the nighttime scene video stream; if the lighting distribution difference exceeds a preset range, then perform adaptive exposure compensation on the real-time video stream and then perform dynamic scene segmentation.

[0012] In some embodiments, performing dynamic scene segmentation on the real-time video stream includes: inputting the real-time video stream into a deep learning model to obtain a semantic segmentation result of the real-time video stream; if there is an unlabeled target category in the semantic segmentation result, marking the target category as a potential abnormal behavior.

[0013] In some embodiments, performing dynamic scene segmentation on the real-time video stream includes: calculating the structural similarity index between adjacent frames in the real-time video stream; if the structural similarity index is lower than a preset threshold, marking the corresponding frame as a dynamic sub-scene area. According to another aspect of the present application, there is provided a security monitoring system supporting AI behavior analysis, including a multi-view camera group and the above-mentioned security monitoring device supporting AI behavior analysis. The security monitoring device and system supporting AI behavior analysis provided by the present application collect a real-time video stream of a target area by using a first camera; perform dynamic scene segmentation on the real-time video stream to obtain multiple sub-scene areas; if it is detected that the behavior pattern of a certain sub-scene area deviates from a preset behavior benchmark, activating a second camera to perform high-resolution supplementary collection on the sub-scene area; wherein, the first camera and the second camera are respectively arranged at different perspective positions of the target area, and the field of view angle of the second camera covers the blind area of the first camera; performing spatio-temporal alignment and fusion on the video streams collected by the first camera and the second camera to generate an enhanced video stream; performing frequency domain transformation on each frame image in the enhanced video stream to extract a behavior feature vector; calculating the similarity difference between the behavior feature vector and a reference behavior model; statistically calculating the maximum and minimum values of the similarity differences of all sub-scene areas within the same time period; determining an abnormal behavior area in the target area based on the maximum and minimum values; that is, first using the first camera to obtain the real-time video stream of the target area and performing dynamic scene segmentation, if an abnormality is found, combining the second camera to perform high-resolution supplementary collection and spatio-temporal alignment and fusion on the video stream, then extracting the behavior feature vector through frequency domain transformation and calculating the maximum and minimum values of the similarity differences, and finally determining the abnormal behavior area, which can not only reduce the computational complexity by relying only on the first camera in a normal scene, but also improve the accuracy of behavior analysis through multi-view fusion in an abnormal scene, thereby ensuring the intelligent level and reliability of the system while improving the detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a schematic structural diagram of a security monitoring device supporting AI behavior analysis provided by an embodiment of the present application, showing the layout relationship between the first camera, the second camera and the target area. Figure 2 It is a schematic flow diagram of dynamic scene segmentation in an embodiment of the present application, including the specific implementation processes of steps such as optical flow vector calculation, edge detection and semantic segmentation. Figure 3This is the processing flow chart for spatio-temporal alignment and fusion of video streams in the embodiments of this application, which details the operation steps of geometric correction, weighted fusion, and denoising processing. Figure 4 This is the logical schematic diagram for determining the abnormal behavior area in the embodiments of this application, showing the processes of similarity difference calculation, maximum and minimum value statistics, and generation of the abnormal behavior heat map. Detailed implementation manners

[0015] This application provides a security monitoring device and system supporting AI behavior analysis. The core lies in achieving efficient detection and precise positioning of abnormal behaviors in the target area through a multi-view camera group and intelligent algorithms. The following combines the attached Figure 1 to the attached Figure 4 and specific embodiments to elaborate in detail the specific implementation manners of the present invention.

[0016] As Figure 1 shown, the security monitoring device in this embodiment includes a first camera 101, a second camera 102, and a target area 103. The first camera 101 is arranged directly above the target area 103 for collecting the real-time video stream of the target area; the second camera 102 is arranged on the side of the target area 103, and its field of view covers the blind area of the first camera 101 to ensure high-resolution supplementary collection of specific sub-scene areas when needed. The background light source of the target area 103 adopts a dimmable LED array 104, which dynamically adjusts the brightness according to the ambient light intensity, thereby providing stable lighting conditions for the collection of the real-time video stream. During actual operation, the first camera 101 first collects the real-time video stream of the target area 103 and transmits the video stream to the processing unit for dynamic scene segmentation.

[0017] The specific process of dynamic scene segmentation is as Figure 2As shown in the figure, it mainly includes steps such as optical flow vector calculation, edge detection, and semantic segmentation. In the optical flow vector calculation, the processing unit extracts the optical flow vector of each pixel point in the real-time video stream frame by frame and analyzes the changes in its direction and magnitude. If the change in the optical flow vector of a certain area exceeds a preset threshold, that area is marked as a dynamic sub-scene area. For example, in a target area at the entrance of a shopping mall, when pedestrians move quickly or vehicles drive in, the trajectory distribution of these moving objects will cause significant changes in the optical flow vector, which are thus identified as dynamic sub-scene areas. In addition, the processing unit also generates an edge image by performing edge detection on the real-time video stream, and further performs connected component analysis on the edge image to obtain multiple connected areas. If the area and shape characteristics of a certain connected area meet the preset conditions (such as an area greater than 500 pixels and an aspect ratio close to 1:2), it is marked as a dynamic sub-scene area. This edge detection-based method is particularly suitable for dynamic scene segmentation in complex backgrounds. For example, in a night scene, due to insufficient light, the optical flow vector calculation may be inaccurate, while edge detection can effectively make up for this defect. To further improve the accuracy of dynamic scene segmentation, this embodiment also introduces a deep learning model.

[0018] As Figure 2 shown, the real-time video stream is input into the deep learning model, and the semantic segmentation result is output. If there is an unlabeled target category in the semantic segmentation result (such as a suddenly appearing unknown object), that target category is marked as a potential abnormal behavior. For example, in a target area of an industrial plant, if the deep learning model detects the appearance of an unlabeled large device in a certain area, it may indicate that there are abnormal operations or safety hazards in that area. After completing the dynamic scene segmentation, the processing unit analyzes the behavior pattern of each sub-scene area. If the behavior pattern of a certain sub-scene area deviates from the preset behavior benchmark, the second camera 102 is activated to perform high-resolution supplementary capture of that sub-scene area. The preset behavior benchmark is a reference behavior model trained based on historical data, and this model is updated distributively through blockchain technology to ensure that it always reflects the latest behavior characteristics. For example, in a target area of a bank hall, under normal circumstances, the walking trajectories of pedestrians are relatively regular. If a large number of people are detected gathering or running quickly during a certain period, it is considered that the behavior pattern of that sub-scene area deviates from the preset behavior benchmark, and at this time, the second camera 102 will be activated to obtain a higher-resolution video stream.

[0019] As Figure 3As shown in the figure, the video streams captured by the first camera 101 and the second camera 102 are then transmitted to the fusion module for spatio-temporal alignment and fusion. In the fusion process, first, geometric correction is performed on the two video streams based on fixed markers in the target area (such as signs on the wall or fixed lines on the ground) to eliminate the perspective deviation caused by different camera positions. Then, the dynamic sub-scene area in the first camera video stream is weighted and fused with the corresponding area in the second camera video stream to obtain a preliminary fusion area. The weighted fusion adopts an adaptive weight assignment strategy, and the weights are determined by the clarity of the data collected by each camera and the trajectory consistency of moving objects. For example, if the data collected by the second camera 102 is clearer and the trajectory consistency of moving objects is stronger, a higher weight is assigned to it. Finally, the preliminary fusion area is denoised to obtain the final fusion area, and an enhanced video stream is generated based on the final fusion area. Each frame image in the enhanced video stream is then sent to the frequency domain transformation module to extract the behavior feature vector. The frequency domain transformation formula is as follows: where F(u, v) is the value at the frequency domain point (u, v), f(x, y) is the grayscale value of the pixel point (x, y), W is the image width, H is the image height, and i is the imaginary unit. This formula converts the spatial domain information into frequency domain information by performing a two-dimensional Fourier transform on the image, thereby extracting the high-frequency components (such as edges and textures) and low-frequency components (such as the background) in the image. These frequency domain information are further transformed into behavior feature vectors for subsequent similarity difference calculation. Calculating the similarity difference between the behavior feature vector and the reference behavior model is one of the key steps in this embodiment. The calculation method of the similarity difference is based on the cosine similarity formula: where A and B are the feature vectors of the behavior feature vector and the reference behavior model respectively, and n is the dimension of the feature vector. By calculating the similarity difference of each sub-scene area, the processing unit statistics the maximum and minimum values of the similarity differences of all sub-scene areas within the same time period.

[0020] Such as Figure 4As shown, if the ratio of the maximum value to the minimum value of the similarity difference within a certain time period exceeds a preset ratio (such as 3:1), the corresponding sub-scene area is marked as an abnormal behavior area. For example, in the target area of a subway station, if the maximum value of the similarity difference in a certain sub-scene area is much higher than that of other areas, it may indicate that there is abnormal behavior in this area, such as crowding or emergencies. To more intuitively display the spatial distribution of abnormal behaviors, this embodiment also generates a heat map of abnormal behaviors in the target area. The depth of the color in the heat map represents the severity of the abnormal behavior, and the darker the color, the higher the possibility of abnormal behavior. For example, in the target area of a large event venue, if the heat map shows that the color of a certain corner is darker, it may indicate that there are potential safety hazards in this area and further investigation is required. In some special scenarios, this embodiment also performs adaptive exposure compensation on the real-time video stream. For example, when switching between the daytime scene and the nighttime scene, the processing unit compares the differences in the light distribution of the target area in the two scenes. If the light distribution difference exceeds the preset range, the real-time video stream is subjected to adaptive exposure compensation before dynamic scene segmentation. This method can effectively cope with the situation of drastic changes in light and ensure the accuracy of dynamic scene segmentation.

[0021] In summary, the security monitoring device and system supporting AI behavior analysis provided by this embodiment collect the real-time video stream of the target area 103 through the first camera 101 and perform dynamic scene segmentation. If an abnormality is detected, high-resolution supplementary capture is performed in combination with the second camera 102, and the video stream is spatially and temporally aligned and fused. Then, the behavior feature vector is extracted through frequency domain transformation, and the maximum and minimum values of the similarity difference are calculated. Finally, the abnormal behavior area is determined. This device can not only reduce the computational complexity by relying only on the first camera in normal scenarios, but also improve the accuracy of behavior analysis through multi-view fusion in abnormal scenarios, thereby ensuring the intelligent level and reliability of the system while improving the detection efficiency.

Claims

1. A security monitoring device supporting AI behavior analysis, characterized in that, Including: A first camera (101) for collecting a real-time video stream of a target area; A processing unit for performing dynamic scene segmentation on the real-time video stream to obtain a plurality of sub-scene areas; wherein, the dynamic scene segmentation is based on the trajectory distribution of moving objects and the background change frequency in the target area; a second camera (102) for performing high-resolution supplementary capture on a certain sub-scene area when its behavior pattern deviates from a preset behavior benchmark, wherein the first camera (101) and the second camera (102) are respectively arranged at different perspective positions of the target area (103), and the field of view angle of the second camera (102) covers the blind area of the first camera (101); a spatio-temporal alignment and fusion module for performing spatio-temporal alignment and fusion on the video streams collected by the first camera (101) and the second camera (102) to generate an enhanced video stream; wherein, the spatio-temporal alignment and fusion adopts an adaptive weight distribution strategy, and the weights are determined by the clarity of the data collected by each camera and the trajectory consistency of moving objects; a frequency domain transformation module for performing frequency domain transformation on each frame image in the enhanced video stream to extract behavior feature vectors; wherein, the frequency domain transformation formula is: wherein, F(u, v) is the value at the frequency domain point (u, v), f(x, y) is the gray value of the pixel point (x, y), W is the image width, H is the image height, and i is the imaginary unit; a behavior analysis module for calculating the similarity difference between the behavior feature vector and a reference behavior model; wherein, the reference behavior model is generated based on historical data training and is updated distributively through blockchain technology; a statistics module for statistically calculating the maximum and minimum values of the similarity differences of all sub-scene areas within the same time period; an abnormal behavior determination module for determining an abnormal behavior area in the target area based on the maximum and minimum values.

2. The security monitoring device supporting AI behavior analysis according to claim 1, wherein The background light source of the target area adopts an adjustable light LED array (104), and the adjustable light LED array dynamically adjusts the brightness according to the environmental light intensity; wherein, performing dynamic scene segmentation on the real-time video stream includes: calculating the optical flow vector of each pixel point in the real-time video stream; if the change in the direction and magnitude of the optical flow vector in a certain area exceeds a preset threshold, then mark this area as a dynamic sub-scene area.

3. The security monitoring device supporting AI behavior analysis according to claim 1, characterized in that, Performing dynamic scene segmentation on the real-time video stream includes: performing edge detection on the real-time video stream to obtain an edge image; performing connected component analysis on the edge image to obtain a plurality of connected regions; if the area and shape characteristics of a certain connected region meet preset conditions, then mark it as a dynamic sub-scene area.

4. The security monitoring device supporting AI behavior analysis according to claim 1, characterized in that, Performing spatio-temporal alignment and fusion on the video streams collected by the first camera (101) and the second camera (102) includes: performing geometric correction on the video streams of the first camera (101) and the second camera (102) based on fixed markers in the target area; performing weighted fusion on the dynamic sub-scene areas in the video stream of the first camera (101) and the corresponding areas in the video stream of the second camera (102) to obtain a preliminary fusion area; performing denoising processing on the preliminary fusion area to obtain a final fusion area; generating the enhanced video stream according to the final fusion area.

5. The security monitoring device supporting AI behavior analysis according to claim 1, wherein Determining the abnormal behavior area in the target area based on the maximum value and the minimum value includes: if the ratio of the maximum value and the minimum value of the similarity difference within a certain time period exceeds a preset ratio, marking the corresponding sub-scene area as an abnormal behavior area; generating an abnormal behavior heat map of the target area according to the spatial distribution of multiple abnormal behavior areas.

6. The security monitoring device supporting AI behavior analysis according to claim 1, characterized in that, The real-time video stream includes the daytime scene video stream and the nighttime scene video stream of the target area; Performing dynamic scene segmentation on the real-time video stream includes: comparing the difference in the light distribution of the target area in the daytime scene video stream and the nighttime scene video stream; If the light distribution difference exceeds a preset range, performing adaptive exposure compensation on the real-time video stream and then performing dynamic scene segmentation.

7. The security monitoring device supporting AI behavior analysis according to claim 1, characterized in that, Performing dynamic scene segmentation on the real-time video stream includes: inputting the real-time video stream into a deep learning model to obtain the semantic segmentation result of the real-time video stream; if there is an unlabeled target category in the semantic segmentation result, marking the target category as a potential abnormal behavior.

8. The security monitoring device supporting AI behavior analysis according to claim 1, characterized in that, Performing dynamic scene segmentation on the real-time video stream includes: calculating the structural similarity index between adjacent frames in the real-time video stream; if the structural similarity index is lower than a preset threshold, marking the corresponding frame as a dynamic sub-scene area.

Citation Information

Patent Citations

  • AI vision-based monitoring method, device, equipment and medium

    CN111814646B

  • An early warning system based on AI video analysis and surveillance security

    CN115457449B