Intelligent security and protection monitoring system and method based on computer vision

By combining monitoring video and sound signals, feature maps are extracted and optimized to generate classified feature vectors, the problem that existing security monitoring systems cannot identify security risks is solved, and real-time and accurate judgment and processing of potential threats are achieved.

CN120408269AInactive Publication Date: 2025-08-01HAINAN JINYUAN VISION COMPUTER TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510485171.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing security monitoring system cannot judge the surveillance scene like manual, and cannot identify potential safety hazards, such as fires or car accidents.

Method used

By fusing monitoring video and sound signals, the area monitoring feature map and sound correlation feature map are extracted, discrete covariant feature entropy optimization is performed, classification feature vectors are generated, and suspicious behavior is judged through the classifier.

Benefits of technology

Real-time and accurate judgment of the monitoring area is achieved, potential security threats can be discovered and handled in a timely manner, and the intelligence level of the monitoring system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408269A_ABST
    Figure CN120408269A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent monitoring, and more specifically discloses an intelligent security and protection monitoring system and method based on computer vision, and the method comprises the steps: obtaining a monitoring video and sound signals of a to-be-detected region, extracting a region monitoring feature map and a region sound correlation feature map, and fusing the two features to obtain a classification feature vector of the to-be-detected region; and a classification result is obtained through a classifier and is used for representing whether a suspicious behavior exists or not.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of intelligent monitoring, and more specifically, to an intelligent security monitoring system and method based on computer vision. Background Art

[0002] A security monitoring system applies optical fibers, coaxial cables or microwaves to transmit video signals within its closed loop, and forms an independent and complete system from camera shooting to image display and recording. It can replace manual work for long-term monitoring in harsh environments, enabling people to see all actual situations occurring at the monitored site and recording them through a video recorder. At the same time, the alarm system equipment alarms for illegal intrusion, and the generated alarm signal is input into the alarm host, which triggers the monitoring system to record and save the video.

[0003] Although the security monitoring system can replace manual work for long-term monitoring, the current security monitoring system cannot judge the monitored scene like a human, such as judging whether there is a fire or a car accident in the scene. This results in the inability to judge things with potential safety hazards, and thus fails to play a warning role.

[0004] Therefore, there is a need for an intelligent security monitoring system and method based on computer vision. Summary of the Invention

[0005] To solve the above technical problems, the present application is proposed. Embodiments of the present application provide an intelligent security monitoring system and method based on computer vision, which realizes the detection and judgment of suspicious behaviors by fusing monitoring videos and sound signals.

[0006] Correspondingly, according to one aspect of the present application, there is provided an intelligent security monitoring system based on computer vision, which includes:

[0007] A detection area data acquisition module for acquiring monitoring videos and sound signals of the area to be detected;

[0008] A detection area data processing module for extracting an area monitoring feature map from the monitoring video and an area sound correlation feature map from the sound signal;

[0009] A detection area data fusion module for performing discrete covariance feature entropy optimization on the area monitoring feature map based on the area sound correlation feature map to obtain a classification feature vector of the area to be detected;

[0010] A detection area data analysis module for passing the classification feature vector of the area to be detected through a classifier to obtain a classification result, where the classification result is used to indicate whether there is a suspicious behavior.

[0011] According to another aspect of the present application, there is also provided an intelligent security monitoring method based on computer vision, which includes:

[0012] Obtaining a monitoring video and a sound signal of the area to be detected;

[0013] Extracting a regional monitoring feature map from the monitoring video, and extracting a regional sound correlation feature map from the sound signal;

[0014] Performing discrete covariance feature entropy optimization on the regional monitoring feature map based on the regional sound correlation feature map to obtain a classification feature vector of the area to be detected;

[0015] Passing the classification feature vector of the area to be detected through a classifier to obtain a classification result, where the classification result is used to indicate whether there is a suspicious behavior.

[0016] Compared with the prior art, an intelligent security monitoring system and method based on computer vision provided by the present application obtain a monitoring video and a sound signal of the area to be detected, extract a regional monitoring feature map and a regional sound correlation feature map, fuse the two to obtain a classification feature vector of the area to be detected, and then obtain a classification result through a classifier, which is used to indicate whether there is a suspicious behavior. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0018] Figure 1 It is a schematic block diagram of an intelligent security monitoring system based on computer vision according to an embodiment of the present application.

[0019] Figure 2 It is a schematic block diagram of a detection area data processing module in an intelligent security monitoring system based on computer vision according to an embodiment of the present application.

[0020] Figure 3 It is a schematic block diagram of a monitoring video processing unit in an intelligent security monitoring system based on computer vision according to an embodiment of the present application.

[0021] Figure 4 It is a schematic block diagram of a sound signal processing unit in an intelligent security monitoring system based on computer vision according to an embodiment of the present application.

[0022] Figure 5It is a flowchart of an intelligent security monitoring method based on computer vision according to an embodiment of the present application. Detailed implementation manners

[0023] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0024] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein is not necessarily to be construed as superior or better than other embodiments.

[0025] In addition, for better illustration of the present application, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present application can also be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present application.

[0026] Furthermore, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality of" means two or more unless otherwise specifically defined.

[0027] Figure 1 The block diagram of an intelligent security monitoring system based on computer vision according to an embodiment of the present application is illustrated. As Figure 1 shown, the intelligent security monitoring system 100 based on computer vision according to an embodiment of the present application includes: a detection area data acquisition module 110 for acquiring a monitoring video and a sound signal of a detection area to be detected; a detection area data processing module 120 for extracting a regional monitoring feature map from the monitoring video and extracting a regional sound correlation feature map from the sound signal; a detection area data fusion module 130 for performing discrete covariant feature entropy optimization on the regional monitoring feature map based on the regional sound correlation feature map to obtain a classification feature vector of the detection area to be detected; and a detection area data analysis module 140 for passing the classification feature vector of the detection area to be detected through a classifier to obtain a classification result, where the classification result is used to indicate whether there is a suspicious behavior.

[0028] In the embodiment of the present application, the detection area data acquisition module 110 is used to obtain the monitoring video and sound signals of the area to be detected. It should be understood that the monitoring video and sound signals provide different perception information. By combining the two, more comprehensive monitoring data can be obtained. Some events may only be recognized through certain features in the sound signal or video signal. Utilizing these two types of information comprehensively can improve the comprehensiveness of monitoring. The video and sound signals can corroborate each other to improve the accuracy of event recognition. Through video monitoring, the occurrence process and details of an event can be seen, while the sound signal can provide additional information, such as human voices, traffic noise, etc., which helps to more accurately judge the type and nature of the event. Combining the monitoring video and sound signals can achieve real-time monitoring of the area to be detected, promptly discover abnormal situations and take corresponding measures. In the field of security monitoring, real-time performance is very important, which can help respond to potential security threats in a timely manner. The comprehensive analysis of video and sound signals can provide more information dimensions, contributing to a deeper understanding of the background and motivation of the event occurrence. Through comprehensive analysis, the authenticity and danger level of the event can be better judged.

[0029] In the embodiment of the present application, the detection area data processing module 120 is used to extract the area monitoring feature map from the monitoring video and the area sound correlation feature map from the sound signal. It should be understood that the monitoring video and sound signals are high-dimensional data, containing a large amount of information and details. Extracting the feature map can transform this complex data into a more concise and representative feature representation, which helps to reduce the dimension and complexity of the data while retaining key information. By extracting the area monitoring feature map and the sound correlation feature map, it can help the computer system identify and understand the patterns and features in the monitoring data. These feature maps can contain information about the appearance, movement, sound characteristics, etc. of the target object, which helps the system perform pattern recognition and target detection. Extracting the feature map enables the data collected at different time points or by different monitoring devices to be subjected to feature matching and comparison. By comparing the similarities and differences between the feature maps, event tracking, recognition, and analysis can be achieved. The feature map is the input data for machine learning and deep learning algorithms, which can help train the model to learn and extract useful information from the monitoring data. In many monitoring systems, using the feature map for machine learning can achieve functions such as intelligent monitoring, behavior recognition, and anomaly detection.

[0030] Specifically, in an embodiment of the present application, Figure 2 The figure shows a block diagram schematic of the detection area data processing module in the intelligent security monitoring system based on computer vision according to the embodiment of the present application. As Figure 2As shown, in the above computer vision-based intelligent security monitoring system 100, the detection area data processing module 120 includes: a monitoring video processing unit 121, configured to extract key frames from the monitoring video and then perform object detection to obtain the area monitoring feature map; and a sound signal processing unit 122, configured to perform noise reduction processing on the sound signal and then perform convolutional coding to obtain the area sound association feature map.

[0031] Correspondingly, in a specific example of the present application, the monitoring video processing unit 121 is configured to extract key frames from the monitoring video and then perform object detection to obtain the area monitoring feature map. It should be understood that a monitoring video is usually a continuous video stream containing a large number of frames and information. By extracting key frames, the computational load and storage space requirements can be reduced, while retaining the most representative and information-rich frames in the video, which is beneficial for subsequent processing and analysis. Object detection technology can help accurately locate key targets or objects in the video and identify specific objects or people in the video frames. Through object detection, the area targets in the monitoring video can be accurately located, providing accurate area information for subsequent feature extraction and analysis. Through object detection, the bounding boxes and position information of the targets can be obtained, and then the features related to the targets can be extracted from the key frames. These features can include information about the appearance, shape, movement, etc. of the targets, which helps to perform a more in-depth analysis and identification of the targets. By extracting key frames and performing object detection, the tracking and tracing of targets in the video sequence can be achieved. This is very important for scenarios in the monitoring system that require real-time monitoring of the target position and behavior, and helps to achieve continuous tracking and analysis of the targets. Through object detection and feature extraction, the recognition accuracy and precision of the monitoring system for targets can be improved. By using object detection and feature extraction technologies, the target information in the monitoring video can be better understood, thus achieving more accurate target recognition and analysis.

[0032] Furthermore, Figure 3 The figure illustrates a block diagram of the monitoring video processing unit in the computer vision-based intelligent security monitoring system according to an embodiment of the present application. As Figure 3 shown, in the monitoring video processing unit 121 of the detection area data processing module 120 of the above computer vision-based intelligent security monitoring system 100, the monitoring video processing unit 121 includes: an extraction area key frame sub-unit 1211, configured to extract multiple area key frames from the monitoring video; an object detection sub-unit 1212, configured to respectively pass the multiple area key frames through an anchor-free object detection network to obtain multiple target object of interest area maps; and a three-dimensional convolution sub-unit 1213, configured to arrange the multiple target object of interest area maps into a three-dimensional tensor and then pass it through a three-dimensional convolutional neural network model to obtain the area monitoring feature map.

[0033] Specifically, the region key frame sub-unit 1211 is used to extract multiple region key frames from the surveillance video. It should be understood that the surveillance video usually contains multiple regions of interest and targets. Extracting multiple region key frames can ensure that the information covering different regions and targets is included, making the analysis more comprehensive and integrated. There may be multiple targets to be tracked and analyzed in the surveillance video. By extracting multiple region key frames, independent feature extraction and processing can be performed for different targets or regions, which is beneficial to realizing multi-target tracking and analysis. The scene in the surveillance video may change, and the key information in different regions may appear at different time points. Extracting multiple region key frames can capture the important information at different time points and scenes, which helps to comprehensively understand the surveillance data. By extracting multiple region key frames, the possibility of missing important information can be reduced. Even if there are no targets in some regions in some frames, other regions can still provide useful information to ensure the integrity of the overall surveillance data. Different regions may provide information from different angles and perspectives. By extracting multiple region key frames, multi-angle analysis can be achieved. This helps to comprehensively understand all aspects of the surveillance scene and helps to improve the understanding and utilization of the surveillance data.

[0034] Correspondingly, the region key frame sub-unit includes: an initial frame extraction secondary sub-unit, which is used to extract an initial frame from the surveillance video; a position difference calculation secondary sub-unit, which is used to calculate the position difference between other frames and the initial frame in the surveillance video to obtain a differential image frame; and a region key frame secondary sub-unit, which is used to determine the multiple region key frames from the surveillance video based on the comparison between the sum value of the feature values of all pixel points in the differential image frame and a predetermined threshold.

[0035] Further, the region key frame secondary sub-unit includes: a first key frame tertiary sub-unit, which is used to determine the other frame corresponding to the differential feature map as the first key frame based on the sum value of the feature values of all pixel points in the differential image frame being greater than the predetermined threshold; and an initial frame setting tertiary sub-unit, which sets the first key frame as the initial frame.

[0036] Specifically, the target detection subunit 1212 is configured to obtain multiple region-of-interest maps of target objects by respectively passing the multiple region key frames through an anchor-free object detection network. It should be understood that the anchor-free object detection network can effectively detect target objects in an image without pre-defining anchor boxes. This network structure can better adapt to targets of different scales, ratios, and shapes, improving the accuracy and robustness of object detection. The multiple region key frames usually contain important regions and targets in the surveillance video. By using the object detection network, the regions of interest of target objects can be accurately identified and located in these key frames, thereby further analyzing the dynamics and behaviors of the targets. The anchor-free object detection network can provide more accurate target localization, identify the bounding boxes of target objects, and accurately describe the position and shape of the targets. This is very important for subsequent target tracking, analysis, and recognition. By obtaining the region-of-interest maps of target objects, the feature information of the targets, such as color, texture, shape, etc., can be further extracted. These features are of great significance for target recognition, classification, and behavior analysis. Combining the multiple region key frames and the region-of-interest maps of target objects can achieve a comprehensive analysis of different regions and targets in the surveillance video. This helps to deeply understand the events and activities occurring in the surveillance scene and improve the utilization value of surveillance data.

[0037] Correspondingly, the target detection subunit includes: a shallow feature Figure 2 level subunit, configured to respectively pass each of the multiple region key frames through multiple convolutional layers to obtain multiple shallow feature maps; an anchor-free detection secondary subunit, configured to respectively pass the multiple shallow feature maps through the anchor-free object detection network to obtain the multiple region-of-interest feature maps of target objects.

[0038] Furthermore, the shallow feature Figure 2The level-0 sub-unit is used to respectively pass each of the multiple regional key frames through multiple convolutional layers to obtain multiple shallow feature maps. It should be understood that multiple convolutional layers are widely used in deep learning networks for feature extraction. Through multiple convolutional operations, different levels and different abstraction levels of features in the image can be gradually extracted, so as to better describe the information in the image. The deep convolutional network has a hierarchical structure from low-level to high-level features, and each layer can learn different-level feature representations. Passing the regional key frames through multiple convolutional layers can extract richer and more abstract feature representations, which helps to better understand and describe the image content. When processing images, the convolutional layer can retain spatial information, and the features at different positions in the image can be captured through the sliding operation of the convolutional kernel. This helps to perform local feature extraction on different regions in the regional key frames, so as to better understand the content and structure within the regions. The convolutional layer has the characteristic of parameter sharing, which can reduce the number of model parameters and the risk of overfitting. This makes the convolutional neural network more efficient and effective in processing image data. The multiple shallow feature maps can be fused and combined through subsequent network structures to obtain a more comprehensive and rich feature representation. This helps to improve the recognition and understanding ability of the target objects in the regional key frames.

[0039] Furthermore, the anchor-free detection secondary sub-unit is used to respectively pass the multiple shallow feature maps through the anchor-free based object detection network to obtain the multiple target object region of interest feature maps. It should be understood that the multiple shallow feature maps contain feature information at different levels and different abstraction levels. By respectively inputting them into the object detection network, the object can be detected and located at different levels, comprehensively using the feature information at each level to improve the accuracy of object detection. Feature maps of different depths can capture information at different scales. Shallow feature maps are more suitable for capturing local detail information, while deep feature maps are more suitable for capturing global semantic information. Through the combination of multiple feature maps, rich multi-scale information can be obtained, which helps to improve the robustness and accuracy of object detection. The context information contained in the multiple feature maps can complement and enrich each other. Through the processing of the object detection network, this context information can be comprehensively utilized to improve the understanding and recognition ability of the target object. The object detection network can learn to reorganize and integrate features at different levels to generate more discriminative feature representations. This helps to accurately locate the target object and extract the region of interest features of the target object. Passing the multiple shallow feature maps through the object detection network for processing can achieve end-to-end training. The network can simultaneously optimize the tasks of feature extraction and object detection through backpropagation, thereby improving the performance and generalization ability of the entire system.

[0040] Specifically, the three-dimensional convolution subunit 1213 is used to arrange the multiple target object region-of-interest maps into a three-dimensional tensor and pass it through a three-dimensional convolutional neural network model to obtain a region monitoring feature map. It should be understood that a three-dimensional convolutional neural network can process data with spatio-temporal dimensions and jointly model the temporal information and spatial information in video data. Arranging the target object region-of-interest maps into a three-dimensional tensor can preserve the information in the time dimension, which helps to better understand the movement trajectories and behaviors of target objects in the surveillance video. A three-dimensional convolutional neural network can effectively extract the feature representations in three-dimensional data. By performing convolutional operations in the spatio-temporal dimensions, it can capture the features at different time points and spatial positions in the video data. This helps to more comprehensively describe the target objects and their surrounding environments in the surveillance video. Arranging the target object region-of-interest maps into a three-dimensional tensor can avoid the information loss that may be introduced when converting temporal information into a planar image. A three-dimensional convolutional neural network can better preserve the temporal information and improve the utilization efficiency of surveillance video data. Through a three-dimensional convolutional neural network, the spatial relationships and dynamic changes between target objects can be learned in three-dimensional space. This helps to better understand the interaction relationships and behavior patterns between target objects in the surveillance video. A three-dimensional convolutional neural network can comprehensively consider the temporal and spatial information and learn richer and more complex feature representations. By arranging the target object region-of-interest maps into a three-dimensional tensor and inputting it into the model, a region monitoring feature map with higher discriminability and representational ability can be obtained..

[0041] Correspondingly, in a specific example of the present application, the sound signal processing unit 122 is configured to perform noise reduction processing on the sound signal and then perform convolutional coding to obtain the regional sound correlation feature map. It should be understood that sound signals are usually interfered by environmental noise. Noise reduction processing can effectively reduce the influence of background noise on the sound signal, improve the quality and distinguishability of the sound signal. This helps the subsequent feature extraction and analysis processes to more accurately capture the key information in the sound signal. Convolutional coding can effectively extract local features and spatial correlations in the signal. Through convolutional operations, the spectral features and time-domain features in the sound signal can be captured. This helps to distinguish the differences and commonalities between different sound signals, thereby better characterizing the characteristics of the sound signal. Convolutional coding can perform feature extraction in the time and frequency dimensions and effectively encode the spatial information in the sound signal. Through convolutional operations, the temporal changes and spectral distributions in the sound signal can be captured, providing a more comprehensive and accurate feature representation for subsequent sound analysis. Passing the noise-reduced sound signal through convolutional coding can fuse the feature information in different frequencies and time domains to generate a richer and more representative sound feature representation. This helps to improve the representation ability and classification accuracy of the sound signal. Convolutional coding can effectively reduce the information redundancy in the signal and retain the key feature information. By performing convolutional coding on the noise-reduced sound signal, the important information in the sound signal can be more accurately extracted and encoded, improving the efficiency and accuracy of sound signal processing.

[0042] Further, Figure 4 The figure shows a block diagram of the sound signal processing unit in an intelligent security monitoring system based on computer vision according to an embodiment of the present application. As Figure 4 shown, in the detection area data processing module 120 of the above-mentioned intelligent security monitoring system 100 based on computer vision, the sound signal processing unit 122 includes: a sound signal noise reduction subunit 1221 configured to perform noise reduction processing on the sound signal to obtain a noise-reduced sound signal; a feature extraction subunit 1222 configured to extract the logarithmic mel spectrogram and cochlear spectrogram of the noise-reduced sound signal based on the area to be detected; a first convolutional coding subunit 1223 configured to pass the logarithmic mel spectrogram through a first convolutional neural network as a filter to obtain a regional sound mel spectrogram feature vector; a second convolutional coding subunit 1224 configured to pass the cochlear spectrogram through a second convolutional neural network as a filter to obtain a regional cochlear spectrogram feature vector; a fusion correlation feature subunit 1225 configured to fuse the regional sound mel spectrogram feature vector and the regional cochlear spectrogram feature vector to obtain a regional sound correlation feature matrix; and a third convolutional coding subunit 1226 configured to pass the regional sound correlation feature matrix through a third convolutional neural network model as a feature extractor to obtain a regional sound correlation feature map.

[0043] Specifically, the sound signal noise reduction subunit 1221 is configured to perform noise reduction processing on the sound signal to obtain a noise-reduced sound signal. It should be understood that the sound signal is usually interfered by noise from the environment, equipment, or the transmission process. Noise reduction processing can effectively reduce these interferences and improve the quality and clarity of the sound signal. Noise will reduce the audibility and recognizability of the sound signal. Noise reduction processing can make the sound signal easier to be recognized and understood by the human auditory system, and improve the clarity and accuracy of the sound signal. In many sound signal processing tasks, noise will affect the feature extraction and analysis process of the signal. Noise reduction processing can help extract the key features in the sound signal, so as to better perform subsequent signal processing and analysis. Performing noise reduction processing on the sound signal can reduce the interference of noise on the sound signal and improve the recognition accuracy and classification performance of the sound signal. This is of great significance for applications such as speech recognition and audio analysis.

[0044] Correspondingly, the feature extraction subunit 1222 is configured to extract the logarithmic mel spectrogram and cochlear spectrogram of the noise-reduced sound signal based on the region to be detected. It should be understood that the logarithmic mel spectrogram and cochlear spectrogram are spectral representations of the sound signal, which can better display the characteristics of the sound signal in the frequency domain. These spectrograms can reflect the intensity and distribution of different frequency components in the sound signal, and are helpful for spectral analysis and feature extraction of the sound signal. The logarithmic mel spectrogram and cochlear spectrogram can extract the frequency features in the sound signal, such as the frequency components and spectral profiles of the sound signal. These frequency features are of great significance for the classification, recognition, and analysis of the sound signal. The logarithmic mel spectrogram and cochlear spectrogram can better characterize the characteristics of the sound signal in the frequency domain and provide a more compact and effective feature representation. These feature representations can be used for tasks such as classification, recognition, and detection of sound signals. The cochlear spectrogram simulates the frequency resolution ability of the human auditory system to the sound signal and can better reflect the processing method of the sound signal in human audition. Through the cochlear spectrogram, the processing process of the sound signal in the auditory system can be better understood. The logarithmic mel spectrogram and cochlear spectrogram are commonly used feature representation methods in the fields of speech recognition and audio processing, which can provide rich frequency feature information and help improve the performance and accuracy of sound signal processing tasks. The logarithmic mel spectrogram and cochlear spectrogram usually have a lower dimension, which can help reduce the dimension of the feature space, reduce the redundancy of features, and thus improve the efficiency and accuracy of sound signal processing.

[0045] Further, the first convolutional encoding subunit 1223 is configured to obtain the regional sound mel spectrogram feature vector by passing the logarithmic mel spectrogram through a first convolutional neural network serving as a filter. It should be understood that a convolutional neural network can effectively extract features from input data and perform abstraction. Through the operations of convolutional layers and pooling layers, a convolutional neural network can learn the spatial features and frequency domain features in the data. With the logarithmic mel spectrogram as the input, the convolutional neural network can learn the spectral features and structural information in the sound signal. A convolutional neural network has a multi-layer structure. By stacking multiple convolutional layers and pooling layers, it can learn the hierarchical feature representation of the data. Inputting the logarithmic mel spectrogram into the first convolutional layer can extract basic spectral features at a lower level, providing a basis for feature extraction and learning in subsequent levels. A convolutional neural network has the characteristics of parameter sharing and weight sharing, which enables the network to effectively utilize the local correlation and spatial structure of the data. After the logarithmic mel spectrogram is processed by the convolutional layer, the network can learn local features, and parameter sharing can reduce the number of parameters in the network and improve the generalization ability of the model. The activation function in the convolutional neural network introduces non-linear mapping, which can enhance the expression ability of the network and better capture the complex relationships in the data. Through non-linear mapping of the logarithmic mel spectrogram in the convolutional layer, the ability of the network to extract and represent the features of the sound signal can be improved. A convolutional neural network can automatically learn the features in the input data without manually designing a feature extractor. Inputting the logarithmic mel spectrogram into the convolutional neural network, the network can automatically learn the optimal feature representation through the backpropagation algorithm, so as to better distinguish different categories of sound signals. Inputting the logarithmic mel spectrogram into the convolutional neural network can achieve end-to-end learning. From the original input data to the final feature representation and classification, the entire process is automatically completed by the network. This simplifies the model design and training process and improves the overall performance of the system.

[0046] Furthermore, the second convolutional encoding subunit 1224 is configured to obtain a regional cochlear spectrogram feature vector by passing the cochlear spectrogram through a second convolutional neural network serving as a filter. It should be understood that the second convolutional neural network can learn more abstract and high-level feature representations. With the cochlear spectrogram as the input, the second convolutional neural network can learn more complex and abstract spectral features, which helps to improve the feature representation ability of the sound signal. The second convolutional neural network can further extract more advanced features based on the basic features learned by the first convolutional neural network. After being processed by the second convolutional neural network, the cochlear spectrogram can extract richer and more abstract spectral features, which helps to improve the feature representation ability of the sound signal. The second convolutional neural network can fuse and combine the features learned from the cochlear spectrogram with the features learned by the first convolutional neural network. This feature fusion and combination can help the network better capture the complex features and structural information of the sound signal and improve the feature representation ability of the sound signal. The deep structure of the second convolutional neural network can help improve the generalization ability of the model, enabling the model to better adapt to different sound signal data. By passing the cochlear spectrogram through the learning of the second convolutional neural network, the generalization ability and robustness of the model to sound signals can be improved. The activation function and convolutional operation in the second convolutional neural network can introduce non-linear mapping, which helps to improve the network's ability to model complex relationships in the data. After being processed by the second convolutional neural network, the cochlear spectrogram can better extract the non-linear features in the sound signal.

[0047] Specifically, the fusion and association feature subunit 1225 is used to fuse the regional sound Mel spectrogram feature vector and the regional cochlear spectrogram feature vector to obtain a regional sound association feature matrix. It should be understood that the regional sound Mel spectrogram feature vector and the regional cochlear spectrogram feature vector respectively capture different aspects and features of the sound signal. Fusing these two different types of feature vectors can comprehensively consider the spectral features and cochlear structure features of the sound signal, improving the diversity and richness of feature representation. The regional sound Mel spectrogram feature vector and the regional cochlear spectrogram feature vector respectively represent the features of the sound signal in the frequency domain and the structural domain. Fusing these two features together can improve the feature representation ability, more comprehensively describe the features of the sound signal, and help improve the classification and recognition performance of the model. The feature information represented by the regional sound Mel spectrogram and the regional cochlear spectrogram is complementary. Fusing these two features can integrate the complementary information between them, make up for the deficiencies of their respective features, and thus improve the richness and robustness of the final feature representation. The regional sound Mel spectrogram and the regional cochlear spectrogram reflect the features of the sound signal in terms of spectrum and structure. Fusing these two features together can comprehensively consider multiple aspects of the sound signal features, making the final feature representation more comprehensive and integrated, and helping to improve the classification and recognition performance of the sound signal. Fusing different types of feature vectors can increase the generalization ability of the model and reduce the risk of overfitting. By comprehensively considering the information of different features, the model can better adapt to different types of data and improve the generalization performance of the model.

[0048] Accordingly, the third convolutional encoding subunit 1226 is configured to obtain a regional sound association feature map by passing the regional sound association feature matrix through a third convolutional neural network model serving as a feature extractor. It should be understood that the third convolutional neural network can learn more abstract and high-level feature representations. By using the regional sound association feature matrix as the input, the third convolutional neural network can further extract and learn more complex and abstract sound association features, which helps to improve the feature representation ability. The third convolutional neural network can further extract more advanced features based on the basic features learned in the first two convolutional layers. By processing the regional sound association feature matrix through the third convolutional neural network, richer and more abstract sound association features can be extracted, which helps to capture the complex features and structural information of the sound signal. The third convolutional neural network can fuse and combine the features learned from the regional sound association feature matrix with the features learned in the first two convolutional layers. This feature fusion and combination can help the network better capture the multi-faceted feature information of the sound signal and improve the representation ability of the sound association features. The deep structure of the third convolutional neural network can help improve the generalization ability of the model, enabling the model to better adapt to different types of sound data. By learning through the third convolutional neural network with the regional sound association feature matrix, the generalization ability and robustness of the model to the sound signal can be improved. The activation function and convolutional operation in the third convolutional neural network can introduce non-linear mapping, which helps to improve the network's ability to model complex relationships in the data. By processing the regional sound association feature matrix through the third convolutional neural network, the non-linear features in the sound signal can be better extracted.

[0049] In the embodiment of the present application, the detection area data fusion module 130 is configured to optimize the discrete covariant feature entropy of the area monitoring feature map based on the area sound association feature map to obtain a classification feature vector of the area to be detected. It should be understood that the area monitoring feature map and the area sound association feature map respectively represent the feature information of the monitoring video and the sound signal. Fusing these two different modalities of information can comprehensively consider the features of the video and the sound, making the feature representation of the area to be detected more comprehensive and rich. The area monitoring feature map and the area sound association feature map contain information in different aspects. By fusing these two types of information, the classification accuracy of the area to be detected can be improved. Combining the features of the video and the sound can provide more information to support accurate classification decisions. Fusing the area monitoring feature map and the area sound association feature map can enhance the ability of feature representation. Video and sound are different perceptual modalities that provide complementary information. Fusing these two types of information together can improve the feature representation ability of the area to be detected, which helps to improve the classification performance. The information contained in the area monitoring feature map and the area sound association feature map is complementary. Fusing these two types of information can integrate the complementary information between them, helping the model to better understand the features of the area to be detected and improving the classification accuracy and robustness. Video and sound are different perceptual modalities that provide different aspects of information about the area to be detected. By fusing these two types of information, multi-faceted information can be comprehensively considered, enabling the model to more comprehensively understand the area to be detected and helping to improve the classification performance.

[0050] Correspondingly, in an embodiment of the present application, the detection area data fusion module 130 is configured to: expand the area monitoring feature map into an area monitoring feature vector, calculate the self-conjugate covariance matrix of the area monitoring feature vector, and perform key dimension analysis on the self-conjugate covariance matrix of the area monitoring feature vector to obtain a set of area monitoring feature base component coding vectors, which is expressed by the formula:

[0051]

[0052] where V represents the area monitoring feature vector, T represents the transpose of the vector, M z represents the self-conjugate covariance matrix, U represents the set of area monitoring feature base component coding vectors, v1, v2, v m respectively represent the first, second, and mth area monitoring feature base component coding vectors, Λ represents the area monitoring feature diagonal matrix after key dimension analysis, and λ1, λ m respectively represent the first and mth eigenvalues on the diagonal of the area monitoring feature diagonal matrix.

[0053] That is, by calculating the self-conjugate covariance matrix to reveal the implicit correlation between the dimensions within the regional monitoring feature vector, key dimension analysis is then used to decouple and orthogonalize the original feature space, stripping away redundant noise and extracting the linear principal components with maximum information entropy. Specifically, the variance maximization criterion is used to screen out a set of regional monitoring feature basic component encoding vectors that can characterize the core dynamic changes of the monitoring scene. This set of regional monitoring feature basic component encoding vectors not only retains the key discriminative information of the original monitoring features, but also improves the sensitivity and generalization ability of the subsequent classifier to abnormal behavior characteristics by eliminating multicollinearity between features.

[0054] In one embodiment of the present application, the detection area data fusion module 130 is further configured to input the set of regional monitoring feature basic component encoding vectors into a sequence encoder based on a forward LSTM model to obtain a set of regional monitoring feature basic component context-related encoding vectors, which is expressed as follows:

[0055] F=LSTM([v1,v2,…,v m ])=[s1,s2,…,s m ]

[0056] Among them, LSTM represents the forward LSTM model, F represents the set of context-related encoding vectors of the basic components of regional monitoring features, s1, s2, s m Represents the context-related encoding vector of the first, second, and mth region monitoring feature basic components.

[0057] That is, by using a forward LSTM to serialize and encode the basic components of the regional monitoring features after PCA sorting, the high-order dynamic association patterns between feature dimensions due to the inherent structure of the data are captured. For example, the temporal memory gating mechanism of the LSTM is used to model the potential causal or complementary relationships between the principal components, thereby transforming the originally isolated static basic components of the regional monitoring features into contextual association encoding vectors of the basic components of the regional monitoring features that contain contextual semantic associations. Specifically, by introducing nonlinear sequence encoding capabilities, while retaining the advantages of orthogonal decoupling of PCA principal components, the implicit hierarchical and structured semantic information in the feature space can be further integrated. This allows the output contextual association encoding vectors of the basic components of the regional monitoring features to simultaneously represent the linear principal component characteristics and nonlinear dynamic association patterns of the monitoring scene, thereby enhancing the subsequent multimodal fusion module's ability to jointly reason about complex abnormal behaviors.

[0058] In one embodiment of the present application, the detection area data fusion module 130 is further configured to: calculate the bitwise rearrangement entropy between each corresponding regional monitoring feature basic component context - associated coding vector and regional monitoring feature basic component coding vector in the set of regional monitoring feature basic component context - associated coding vectors and the set of regional monitoring feature basic component coding vectors to obtain a set of bitwise rearrangement entropies, which is expressed by the formula:

[0059]

[0060] where w represents a bitwise comparison function, v i represents the i - th regional monitoring feature basic component coding vector, represents the feature value at the j - th position of the i - th regional monitoring feature basic component coding vector, s i represents the i - th regional monitoring feature basic component context - associated coding vector, represents the feature value at the j - th position of the i - th regional monitoring feature basic component context - associated coding vector, ε represents a predetermined threshold, r i represents the i - th bitwise displacement comparison feature vector, represents the feature value at the j - th position of the i - th bitwise displacement comparison feature vector, L represents the length of the i - th bitwise displacement comparison feature vector, e i represents the i - th bitwise rearrangement entropy.

[0061] That is, introducing the bitwise rearrangement entropy can start from the perspective of information theory and analyze the microscopic perturbation of the information structure of features before and after coding through the binary bit change pattern. For example, it can capture the flipping or displacement of specific bits in the binary representation of PCA principal components by LSTM sequence coding, so as to more sensitively quantify the information gain or reconstruction strength brought by context - associated modeling. In this way, by comparing the binary - bit - level differences between the regional monitoring feature basic component context - associated coding vector and the regional monitoring feature basic component coding vector, the dynamic change trajectory of information entropy in the feature optimization process can be analyzed. For example, it can identify which bit changes dominate the migration or enhancement of feature semantics, and convert this bit - level perturbation pattern into a quantifiable entropy value signal, providing a refined difference basis at the information - theory level for weight assignment and feature selection in subsequent multi - modal fusion.

[0062] In one embodiment of the present application, the detection area data fusion module 130 is further configured to: perform a weighting process on the set of bitwise rearrangement entropies to obtain a set of bitwise rearrangement entropy adjustment coefficients, which is expressed by the formula:

[0063] a i = softmax(e i )

[0064] Among them, softmax represents the normalized exponential function, and a i represents the i-th bit-by-bit rearrangement entropy adjustment coefficient.

[0065] That is, by using the exponential amplification effect and probability normalization characteristics of Softmax, the information reconstruction intensity characterized by the entropy value of the bit-by-bit rearrangement entropy is transformed into a weight ratio with clear physical meaning. It not only strengthens the focusing ability on key hidden danger features through the steepening of the weight distribution, but also avoids the optimization oscillation caused by extreme weight allocation through probability smoothness. Finally, a bit-by-bit rearrangement entropy adjustment coefficient driven by information-theoretic difference is generated to improve the sensitivity and classification robustness of the system to concealed abnormal behaviors.

[0066] In an embodiment of the present application, the detection area data fusion module 130 is further configured to: based on the set of the bit-by-bit rearrangement entropy adjustment coefficients, fuse the set of the region monitoring feature base component coding vectors to obtain a classification feature vector of the area to be detected, which is expressed by the formula:

[0067]

[0068] where v f represents the classification feature vector of the area to be detected.

[0069] That is, through the strong correlation between the information perturbation intensity and the classification task requirements, a classification feature vector of the area to be detected that retains the orthogonality of the PCA principal components and fuses the context relevance of the LSTM is generated, so that the spatio-temporal features of fire spread and the high-frequency features of abnormal sounds form a cross-modal strong coupling under the guidance of the binary bit-shift entropy. Finally, the capture ability and discrimination accuracy of the classifier for concealed safety hazards are improved to meet the requirements of the security monitoring system for refined identification of abnormal behaviors in complex scenarios.

[0070] In an embodiment of the present application, the detection area data analysis module 140 is configured to obtain a classification result by passing the classification feature vector of the area to be detected through a classifier, and the classification result is used to indicate whether there is a suspicious behavior. It should be understood that in order to detect and identify suspicious behaviors, it is necessary to classify the area to be detected to determine whether the behavior occurring in the area is a normal behavior or a suspicious behavior. The classifier can output a corresponding classification result according to the classification feature vector of the area to be detected, thereby helping to distinguish different types of behaviors. The classifier can help identify behaviors that are different from normal behaviors and may be suspicious or abnormal. By inputting the classification feature vector of the area to be detected into the classifier, it can be determined whether there are features of possible suspicious behaviors in the area, so as to perform further analysis and processing. The result output by the classifier can be used as a basis for decision support to help the system determine whether it is necessary to trigger further behavior recognition or alarm mechanisms. Using the classification result to indicate whether there is a suspicious behavior can help the system promptly discover potential security problems. Applying the classifier to the area to be detected can achieve automatic analysis and processing of the surveillance video. The classification result can intuitively indicate whether there is a suspicious behavior in the area to be detected, thereby reducing the burden of manual monitoring and improving the efficiency of the monitoring system.

[0071] Correspondingly, in an embodiment of the present application, the detection area data analysis module 140 is configured to: use the classifier to process the classification feature vector of the area to be detected according to the following formula to obtain the classification result;

[0072] wherein, the formula is: O = softmax{(W c , B c )|X}, where X represents the classification feature vector of the area to be detected, W c is the weight matrix, B c represents the bias vector, softmax represents the normalized exponential function, and O represents the classification result.

[0073] In summary, based on the computer vision-based intelligent security monitoring system and method according to the embodiments of the present application, by acquiring the surveillance video and sound signal of the area to be detected, extracting the area surveillance feature map and the area sound association feature map, fusing the two to obtain the classification feature vector of the area to be detected, and then obtaining the classification result through the classifier, which is used to indicate whether there is a suspicious behavior.

[0074] As described above, the computer vision-based intelligent security monitoring system 100 according to the embodiments of the present application can be implemented in various terminal devices, such as the server of the computer vision-based intelligent security monitoring system. In one example, the computer vision-based intelligent security monitoring system 100 can be integrated into the terminal device as a software module and / or a hardware module. For example, the computer vision-based intelligent security monitoring system 100 can be a software module in the operating system of the terminal device, or can be an application program developed for the terminal device; of course, the computer vision-based intelligent security monitoring system 100 can also be one of the many hardware modules of the terminal device.

[0075] Alternatively, in another example, the computer vision-based intelligent security monitoring system 100 and the terminal device can also be separate devices, and the computer vision-based intelligent security monitoring system 100 can be connected to the terminal device through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.

[0076] Figure 5 FIG. is a flowchart of a computer vision-based intelligent security monitoring method according to an embodiment of the present application. As Figure 5 shown, the computer vision-based intelligent security monitoring method according to the embodiments of the present application includes the steps of: S110, obtaining a monitoring video and a sound signal of a region to be detected; S120, extracting a region monitoring feature map from the monitoring video and extracting a region sound correlation feature map from the sound signal; S130, performing discrete covariance feature entropy optimization on the region monitoring feature map based on the region sound correlation feature map to obtain a region to be detected classification feature vector; S140, passing the region to be detected classification feature vector through a classifier to obtain a classification result, and the classification result is used to indicate whether there is a suspicious behavior.

[0077] Here, those skilled in the art can understand that the specific operations of each step in the above computer vision-based intelligent security monitoring method have been described in detail in the description of the computer vision-based intelligent security monitoring system above, and therefore, the repeated description thereof will be omitted. Figures 1 to 4 The detailed description of the computer vision-based intelligent security monitoring system has been given above, and thus, the repeated description thereof will be omitted.

[0078] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0079] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0080] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0081] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0082] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0083] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0084] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the said claims.

Claims

1. An intelligent security monitoring system based on computer vision, characterized in that, Including: A detection area data acquisition module, configured to obtain a monitoring video and a sound signal of an area to be detected; A detection area data processing module, configured to extract an area monitoring feature map from the monitoring video and extract an area sound correlation feature map from the sound signal; A detection area data fusion module, configured to perform discrete covariant feature entropy optimization on the area monitoring feature map based on the area sound correlation feature map to obtain a classification feature vector of the area to be detected; A detection area data analysis module, configured to pass the classification feature vector of the area to be detected through a classifier to obtain a classification result, where the classification result is used to indicate whether there is a suspicious behavior.

2. The intelligent security monitoring system based on computer vision according to claim 1, wherein The detection area data processing module includes: A monitoring video processing unit, configured to extract key frames from the monitoring video and then perform object detection to obtain the area monitoring feature map; A sound signal processing unit, configured to perform noise reduction processing on the sound signal and then perform convolutional coding to obtain the area sound correlation feature map.

3. The intelligent security monitoring system based on computer vision according to claim 2, wherein The monitoring video processing unit includes: An area key frame extraction subunit, configured to extract a plurality of area key frames from the monitoring video; An object detection subunit, configured to respectively pass the plurality of area key frames through an anchor-free object detection network to obtain a plurality of region of interest maps of target objects; A three-dimensional convolution subunit, configured to arrange the plurality of region of interest maps of target objects into a three-dimensional tensor and then pass it through a three-dimensional convolutional neural network model to obtain an area monitoring feature map.

4. The intelligent security monitoring system based on computer vision according to claim 3, characterized in that, The area key frame extraction subunit includes: An initial frame extraction secondary subunit, configured to extract an initial frame from the monitoring video; A position-based difference secondary subunit, configured to calculate the position-based difference between other frames in the monitoring video and the initial frame to obtain a difference image frame; An area key frame secondary subunit, configured to determine the plurality of area key frames from the monitoring video based on the comparison between the sum value of the feature values of all pixel points in the difference image frame and a predetermined threshold.

5. The intelligent security monitoring system based on computer vision according to claim 4, characterized in that, The area key frame secondary subunit includes: A first key frame tertiary subunit, configured to determine an other frame corresponding to the difference feature map as a first key frame based on the sum value of the feature values of all pixel points in the difference image frame being greater than a predetermined threshold; An initial frame setting tertiary subunit, configured to set the first key frame as the initial frame.

6. The intelligent security monitoring system based on computer vision according to claim 5, characterized in that, The object detection subunit includes: A shallow feature map secondary subunit, configured to respectively pass each area key frame in the plurality of area key frames through a multi-layer convolutional layer to obtain a plurality of shallow feature maps; An anchor-free detection secondary subunit, configured to respectively pass the plurality of shallow feature maps through the anchor-free object detection network to obtain the plurality of region of interest feature maps of target objects.

7. The intelligent security monitoring system based on computer vision according to claim 6, characterized in that The sound signal processing unit includes: A sound signal noise reduction subunit, configured to perform noise reduction processing on the sound signal to obtain a noise-reduced sound signal; A feature extraction subunit, configured to extract a logarithmic mel spectrogram and a cochlear spectrogram of the noise-reduced sound signal based on the area to be detected; The first convolutional encoding subunit is used to obtain a regional sound mel spectrogram feature vector by passing the logarithmic mel spectrogram through a first convolutional neural network serving as a filter; The second convolutional encoding subunit is used to obtain a regional cochlear spectrogram feature vector by passing the cochlear spectrogram through a second convolutional neural network serving as a filter; The fusion and correlation feature subunit is used to fuse the regional sound mel spectrogram feature vector and the regional cochlear spectrogram feature vector to obtain a regional sound correlation feature matrix; The third convolutional encoding subunit is used to obtain a regional sound correlation feature map by passing the regional sound correlation feature matrix through a third convolutional neural network model serving as a feature extractor.

8. The computer vision-based intelligent security monitoring system according to claim 7, wherein, The detection area data fusion module is used for: Unfolding the regional monitoring feature map into a regional monitoring feature vector, calculating the self-conjugate covariance matrix of the regional monitoring feature vector, and performing critical dimension analysis on the self-conjugate covariance matrix of the regional monitoring feature vector to obtain a set of regional monitoring feature base component coding vectors; Inputting the set of regional monitoring feature base component coding vectors into a sequence encoder based on a forward LSTM model to obtain a set of regional monitoring feature base component context correlation coding vectors; Calculating the bit-by-bit rearrangement entropy between each pair of corresponding regional monitoring feature base component context correlation coding vectors and regional monitoring feature base component coding vectors in the set of regional monitoring feature base component context correlation coding vectors and the set of regional monitoring feature base component coding vectors to obtain a set of bit-by-bit rearrangement entropies; Performing a weighting process on the set of bit-by-bit rearrangement entropies to obtain a set of bit-by-bit rearrangement entropy adjustment coefficients; Based on the set of bit-by-bit rearrangement entropy adjustment coefficients, fusing the set of regional monitoring feature base component coding vectors to obtain a classification feature vector for the area to be detected.

9. The intelligent security monitoring system based on computer vision according to claim 8, characterized in that, The detection area data analysis module is used for: processing the classification feature vector for the area to be detected using the classifier according to the following formula to obtain the classification result; Among them, the formula is: O = softmax{(W c , B c )|X}, where X represents the classification feature vector of the area to be detected, W c is the weight matrix, B c represents the bias vector, softmax represents the normalized exponential function, and O represents the classification result.

10. An intelligent security monitoring method based on computer vision, characterized in that, Including: Obtaining the monitoring video and sound signal of the area to be detected; Extracting a regional monitoring feature map from the monitoring video and a regional sound correlation feature map from the sound signal; Performing discrete covariance feature entropy optimization on the regional monitoring feature map based on the regional sound correlation feature map to obtain a classification feature vector for the area to be detected; Passing the classification feature vector for the area to be detected through a classifier to obtain a classification result, where the classification result is used to indicate whether there is a suspicious behavior.

Citation Information

Patent Citations

  • Mining mechanical equipment anomaly detection system and method

    CN120314678A

  • Intelligent logistics distribution system and method based on Internet of Vehicles system

    CN120373989A

  • Indoor decoration intelligent analysis system and method based on user demands and spatial data

    CN120410589A

  • Enterprise financial information intelligent analysis system and method based on big data

    CN120410740A

  • Agricultural machine remote monitoring system based on big data and block chain technology

    CN120410765A

Cited By

  • Mining mechanical equipment anomaly detection system and method

    CN120314678A

  • Intelligent logistics distribution system and method based on Internet of Vehicles system

    CN120373989A

  • Intelligent customer service implementation system and method based on digital human

    CN120410543A

  • Indoor decoration intelligent analysis system and method based on user demands and spatial data

    CN120410589A

  • Agricultural machine remote monitoring system based on big data and block chain technology

    CN120410765A