Security camera abnormal behavior identification method and system based on multi-modal fusion

By performing feature extraction and weighted fusion processing on the multimodal data collected by security cameras and dynamically adjusting the fusion weights, the problem of insufficient environmental adaptability in existing technologies is solved, and efficient abnormal behavior recognition and alarm are achieved in complex environments.

CN120808275APending Publication Date: 2025-10-17SHENZHEN KEAN DIGITAL CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202511218119.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The existing multimodal data fusion method in security cameras cannot dynamically adjust the weight of each modal data according to the actual environment, resulting in insufficient accuracy and reliability in identifying abnormal behavior in complex environments, affecting the overall effectiveness of the security monitoring system.

Method used

By obtaining the visible light image data, infrared image data and audio signal data collected by the security camera, and performing feature extraction processing on them respectively, a feature vector set is generated, and the weighted fusion algorithm is used to dynamically adjust the fusion weight according to the confidence weight of each modal feature vector. Finally, it is input into the abnormal behavior classification model for classification to generate an abnormal behavior alarm signal.

Benefits of technology

It improves the accuracy and effectiveness of abnormal behavior identification in complex environments, enhances the adaptability and stability of the security monitoring system, and can detect and alarm abnormal behavior in a timely manner to ensure the safety of the monitored area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808275A_ABST
    Figure CN120808275A_ABST
Patent Text Reader

Abstract

The invention relates to a security camera abnormal behavior identification method and system based on multi-modal fusion. The method comprises the steps of obtaining multi-modal data such as visible light image data, infrared image data and audio signal data of a security camera in a target monitoring period; performing feature extraction on the data of different modes to obtain respective feature vector sets; the features of all the modes are fused, and in the multi-mode feature fusion process, a weighted fusion algorithm for dynamically adjusting the fusion weight according to the confidence coefficient weight of feature vectors of all the modes is adopted; inputting the fusion feature vector set into an abnormal behavior classification model for classification and early warning; according to the scheme, the accuracy and effectiveness of abnormal behavior recognition in a complex environment can be improved, and then the overall efficiency of security monitoring is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of security and artificial intelligence, in particular to a security camera abnormal behavior recognition method and system based on multi-modal fusion. BACKGROUND

[0002] In the field of security monitoring, accurate identification of abnormal behavior is a key link to protect public safety and property safety. With the development of technology, security cameras have become an important tool for monitoring, and the types of data they collect have become increasingly diverse, including visible light images, infrared images, and audio signals. The use of multi-modal data makes it possible to more comprehensively and accurately identify abnormal behavior, as it can capture feature information from different angles, providing more information and complementarity than single-modal data.

[0003] To utilize multi-modal data for abnormal behavior recognition, some existing methods use simple data fusion methods, such as directly concatenating or averaging different modal data. These methods improve the accuracy of recognition to some extent, but do not fully consider the reliability and importance differences of each modal data in different scenarios.

[0004] However, these solutions have obvious limitations. Simple data fusion methods cannot dynamically adjust the weights of each modal data according to the actual environment, so that in some cases with heavy environmental interference, some less reliable data may have a negative impact on the recognition result, making it difficult to accurately identify complex and variable abnormal behavior, thereby reducing the overall effectiveness of the security monitoring system. SUMMARY

[0005] The main purpose of the present application is to provide a security camera abnormal behavior recognition method and system based on multi-modal fusion, which can improve the accuracy and effectiveness of abnormal behavior recognition in complex environments, and thus improve the overall effectiveness of security monitoring.

[0006] To achieve the above purpose, the embodiment of the present application provides a security camera abnormal behavior recognition method based on multi-modal fusion, which comprises: acquiring multi-modal data collected by a security camera in a target monitoring period, the multi-modal data including visible light image data, infrared image data, and audio signal data; performing first feature extraction processing on the visible light image data to generate a first feature vector set, performing second feature extraction processing on the infrared image data to generate a second feature vector set, and performing third feature extraction processing on the audio signal data to generate a third feature vector set; inputting the first feature vector set, the second feature vector set and the third feature vector set into a multi-modal feature fusion module, and generating a fusion feature vector set through a weighted fusion algorithm, wherein the weighted fusion algorithm dynamically adjusts fusion weights according to confidence weights of the modal feature vectors; inputting the fusion feature vector set into an abnormal behavior classification model to generate an abnormal behavior classification result, the abnormal behavior classification result including a normal behavior label and an abnormal behavior label; According to the abnormal behavior classification result, an abnormal behavior alarm signal is generated and sent to a security monitoring terminal to trigger an alarm operation.

[0007] In summary, by using the technical solution of the present application, the multi-modal data such as visible light image data, infrared image data and audio signal data of the security camera in the target monitoring period can be obtained, which can comprehensively describe the behavior characteristics in the monitoring scene from multiple dimensions. Feature extraction is performed on different modal data to obtain respective feature vector sets, which provides rich feature information for subsequent fusion processing. In the multi-modal feature fusion process, the weighted fusion algorithm dynamically adjusts the fusion weights according to the confidence weights of the modal feature vectors, which can fully consider the reliability difference of different modal data in different environments, and improves the accuracy and effectiveness of the fusion feature vector set. Inputting the fusion feature vector set into the abnormal behavior classification model for classification can improve the accuracy and effectiveness of abnormal behavior recognition in complex environments, and further improve the overall efficiency of security monitoring.

[0008] Further, in the training process of the abnormal behavior classification model, by obtaining the training data set, performing cross-validation and adjusting the hyperparameter configuration, the classification performance and generalization ability of the model are improved. At the same time, the training data set is preprocessed to ensure the data quality of the input model, further improving the accuracy of the model. In addition, the weighted fusion algorithm of the multi-modal feature fusion module also has a mechanism for dynamically adjusting the confidence weight value, which can adapt to the environmental changes in different monitoring periods, enhancing the adaptability and stability of the system in complex environments. Finally, according to the abnormal behavior classification result, an alarm signal is generated and sent to the security monitoring terminal, which can timely discover and warn abnormal behavior, effectively protecting the safety of the security monitoring area. BRIEF DESCRIPTION OF DRAWINGS

[0009] Figure 1 is a scene diagram of the abnormal behavior recognition method of the security camera based on multi-modal fusion in the embodiment of the present application; Figure 2 is a flowchart of the abnormal behavior recognition method of the security camera based on multi-modal fusion provided by the embodiment of the present application; Figure 3A flowchart of the process of calculating the fusion feature vector set provided by the embodiment of the present application is shown in the figure; Figure 4 Another flowchart of the process of calculating the confidence weight provided by the embodiment of the present application is shown in the figure; Figure 5 A flowchart of the process of model training provided by the embodiment of the present application is shown in the figure; Figure 6 A flowchart of the process of model evaluation provided by the embodiment of the present application is shown in the figure; Figure 7 A flowchart of the process of dynamic adjustment of the confidence weight provided by the embodiment of the present application is shown in the figure; Figure 8 A flowchart of the process of calculating the environmental interference factor value provided by the embodiment of the present application is shown in the figure; Figure 9 A flowchart of the process of calculating the comprehensive environmental parameter provided by the embodiment of the present application is shown in the figure; Figure 10 A structural diagram of the security camera abnormal behavior recognition system based on multi-modal fusion provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0010] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0011] The present application provides a security camera abnormal behavior recognition method and system based on multi-modal fusion, which will be described in detail below.

[0012] In the embodiments of the present application, the security camera abnormal behavior recognition method based on multi-modal fusion is a comprehensive security monitoring technology. It uses multi-modal data such as visible light image data, infrared image data and audio signal data collected by the security camera, and realizes accurate judgment of abnormal behavior through feature extraction, fusion processing and classification recognition of different modal data. Among them, multi-modal data can reflect the behavior characteristics in the monitoring scene from different angles, with the characteristics of high information richness and strong complementarity. Feature extraction processing converts each modal data into a feature vector set, which is convenient for subsequent analysis and processing. The multi-modal feature fusion module uses a weighted fusion algorithm to dynamically adjust the fusion weight according to the confidence weight of each modal feature vector, and generates a fusion feature vector set to improve the accuracy and effectiveness of feature information. The abnormal behavior classification model classifies the fusion feature vector set and outputs normal behavior labels and abnormal behavior labels. The whole process is dynamic and adaptive, which can accurately identify abnormal behavior in different environments and provide reliable technical support for security monitoring As shown in Figure 1 A security camera abnormal behavior recognition method based on multi-modal fusion is provided. In the scene, it mainly includes a security camera, a data processing server and a security monitoring terminal; wherein the security camera, the data processing server and the security monitoring terminal are connected through a network.

[0013] Taking a large shopping mall scene as an example, in this scene, personnel flow frequently, various behavior activities are complex and diverse, and the importance of security monitoring is self-evident. Among them, the security camera is distributed in every corner of the mall, including the entrance, the corridor, the interior of the store and other positions, which is used to collect multi-modal data in different areas.

[0014] Among them, the visible light image data can clearly record the appearance characteristics, action posture and behavior trajectory of the personnel and other information. For example, in the corridor of the mall, the visible light camera can capture the walking direction of the personnel, whether there is running, pushing and other abnormal actions. The infrared image data can reflect the body temperature distribution and thermal radiation characteristics of the personnel, which plays an important role in monitoring the activities of personnel in the night or dark environment. For example, in the warehouse area of the mall, the infrared camera can detect whether there is abnormal heat source activity, such as whether someone enters the warehouse during non-working hours. The audio signal data can capture the sound information on the scene, such as shouting, fighting and other sounds, which provides additional clues for the identification of abnormal behavior. For example, in the catering area of the mall, the audio camera can monitor whether there is a sound of quarrel or conflict.

[0015] After obtaining the multi-modal data uploaded by the security camera, the data processing server performs feature extraction and fusion processing on the data. First, the visible light image data is subjected to first feature extraction processing to generate a first feature vector set. Convolutional neural network (CNN) and other methods can be used to extract features such as texture, shape, and color from the image. For infrared image data, second feature extraction processing is performed to generate a second feature vector set. Heat imaging feature extraction algorithms can be used to extract features such as temperature distribution and hotspot shape from the infrared image. The audio signal data is subjected to third feature extraction processing to generate a third feature vector set. Frequency spectrum analysis and other methods can be used to extract features such as frequency and intensity of the audio signal.

[0016] Then, the first feature vector set, the second feature vector set, and the third feature vector set are input into a multi-modal feature fusion module to generate a fused feature vector set through a weighted fusion algorithm. In this process, the fusion weights are dynamically adjusted according to the confidence weights of the modal feature vectors to fully exploit the advantages of each modal data. For example, in the case of good light during the day, the confidence weight of the visible light image data may be higher; while in the night or in an environment with a lot of smoke, the confidence weight of the infrared image data may be correspondingly increased.

[0017] Next, the fused feature vector set is input into an abnormal behavior classification model for classification. The abnormal behavior classification model is trained and can output normal behavior labels and abnormal behavior labels according to the fused feature vector set. For example, if a person is detected running, fighting, or stealing in a shopping mall, the model will output an abnormal behavior label.

[0018] Finally, the security monitoring terminal generates an abnormal behavior alarm signal according to the abnormal behavior classification result and triggers an alarm operation. For example, when receiving the abnormal behavior alarm signal, the monitoring terminal will issue an audible and visual alarm, and display the location and related image information of the abnormal behavior on the monitoring screen, so that security personnel can take timely measures. At the same time, the system can also store and analyze the related data of the abnormal behavior to provide a reference for subsequent security management.

[0019] Reference Figure 2 , Figure 2 is a flowchart of an abnormal behavior recognition method for a security camera based on multi-modal fusion provided by an embodiment of the present application. The execution subject of the method can be a computer device, which can be a computer device or a cluster composed of multiple computer devices. The computer device can be a terminal device or a server, etc. The abnormal behavior recognition method for a security camera based on multi-modal fusion provided by an embodiment of the present application specifically includes: S10: Obtain multi-modal data collected by the security camera in a target monitoring period, the multi-modal data including visible light image data, infrared image data, and audio signal data.

[0020] In this application, the target monitoring period is a pre-set time range, and in this time range, the security camera continuously collects multi-modal data. Visible light image data is image information captured by the visible light sensor of the camera, which can present the visual features of the monitoring scene, such as the shape, color, and position of objects, etc. For example, in a monitoring scene of an office building, visible light image data can clearly show the dress, posture, and placement of office equipment. Visible light image data plays an important role in identifying facial features, behavior, and object types in the scene, and can provide intuitive visual information for abnormal behavior recognition. In terms of technical effect, accurate acquisition of visible light image data can help identify normal and abnormal behavior patterns of personnel, such as whether a stranger enters a sensitive area, whether there is a violation of operation, etc.

[0021] Among them, the infrared image data is collected by the infrared sensor of the camera, which reflects the thermal radiation information of the object surface. In different environmental temperatures, the thermal radiation characteristics of the object will be different, and by analyzing the infrared image data, information hidden in the visible light image can be detected. For example, in the night or low light environment, infrared image data can clearly show the body temperature distribution and activity track of personnel. Infrared image data is of great significance for detecting the presence of personnel, identifying heat sources, and discovering potential safety hazards. Its technical effect lies in supplementing the deficiency of visible light image data in special environments, improving the accuracy and reliability of abnormal behavior recognition.

[0022] Audio signal data is sound information recorded by the audio sensor equipped with the camera, which contains various sound characteristics in the monitoring scene, such as human voice, object collision sound, alarm sound, etc. For example, in a public place, audio signal data can capture the sound of people talking, shouting, and possible abnormal sound. Audio signal data can provide additional clues for abnormal behavior recognition from the perspective of sound, complementing visible light image data and infrared image data. From the technical effect, audio signal data can help identify some abnormal behaviors that are difficult to judge by image alone, such as arguing, fighting, etc.

[0023] By obtaining these three types of multi-modal data, the embodiments of the present application can comprehensively and multi-angle describe the behavior characteristics in the monitoring scene, providing rich information sources for subsequent abnormal behavior recognition. From the technical effect, the comprehensive use of multi-modal data can improve the accuracy and reliability of abnormal behavior recognition, and reduce the occurrence of misjudgment and omission.

[0024] In an embodiment, in order to ensure that high-quality multi-modal data is obtained, regular maintenance and calibration of the security camera can be performed. For visible light image data, the focal length, aperture, and exposure time of the camera can be adjusted to ensure that the image clarity and brightness are appropriate. At the same time, a high-resolution visible light sensor is used to improve the image detail capture capability. For infrared image data, the infrared sensor can be temperature calibrated to ensure that it can accurately reflect the thermal radiation information of the object. In addition, an infrared sensor with a wide dynamic range is selected to adapt to the monitoring needs under different environmental temperatures. For audio signal data, a high-quality audio sensor can be installed and sound calibrated to improve the quality of sound collection. At the same time, noise reduction technology is used to reduce the interference of environmental noise on the audio signal. Through these measures, the quality of multi-modal data can be improved to provide a good foundation for subsequent processing and analysis.

[0025] S20: performing first feature extraction processing on the visible light image data to generate a first feature vector set, performing second feature extraction processing on the infrared image data to generate a second feature vector set, and performing third feature extraction processing on the audio signal data to generate a third feature vector set.

[0026] In this application, the first feature extraction processing is a series of operations performed on the visible light image data, aiming to extract information that can represent the characteristics of the image from the image. Various feature extraction algorithms can be used, such as convolutional neural network (CNN), local binary pattern (LBP), etc. Convolutional neural network is a deep learning model that automatically learns features in images through convolutional layers, pooling layers, and fully connected layers, etc. For example, in a face recognition scenario, convolutional neural network can learn the features of face contours, eyes, nose, etc. Local binary pattern is a feature extraction method based on local texture of images, which generates a binary pattern by comparing the gray values of each pixel and its neighborhood pixels in the image, to describe the texture features of the image. The first feature vector set is a vector set composed of extracted features, which can represent the feature information of the visible light image data. Its technical effect lies in converting complex image data into feature vectors that are easy to process and analyze, improving the efficiency and accuracy of subsequent processing.

[0027] The second feature extraction processing is a feature extraction operation on the infrared image data. Since the infrared image mainly reflects the thermal radiation information of the object, some feature extraction methods for thermal images can be used, such as thermal gradient feature extraction, thermal texture feature extraction, etc. The thermal gradient feature extraction reflects the temperature change of the object surface by calculating the thermal gradient value of the pixels in the infrared image. For example, in an industrial equipment monitoring scene, the thermal gradient feature can detect whether the equipment has overheating or abnormal temperature distribution. The thermal texture feature extraction is to extract features from the texture of the infrared image, which is used to describe the thermal distribution pattern of the object surface. The second feature vector set is a vector set composed of the extracted infrared image features, which can effectively represent the feature information of the infrared image data. From the technical effect, by extracting features from the infrared image data, the hidden information can be mined, and more basis can be provided for abnormal behavior recognition.

[0028] The third feature extraction processing is a feature extraction operation on the audio signal data. Some common audio feature extraction methods can be used, such as Mel-frequency cepstral coefficient (MFCC), linear predictive cepstral coefficient (LPCC), etc. Mel-frequency cepstral coefficient is a widely used feature extraction method in speech recognition, which extracts coefficients representing the characteristics of audio signals through spectral analysis and cepstral transform of audio signals. For example, in a speech recognition system, Mel-frequency cepstral coefficients can be used to distinguish the speech characteristics of different people. Linear predictive cepstral coefficient is a method based on linear prediction analysis, which extracts the cepstral coefficients of prediction errors by linear prediction of audio signals, and is used to describe the spectral characteristics of audio signals. The third feature vector set is a vector set composed of the extracted audio features, which can accurately represent the feature information of the audio signal data. Its technical effect is to convert the audio signal into a feature vector, which is convenient for subsequent classification and recognition processing.

[0029] The embodiments of the present application generate respective feature vector sets by performing feature extraction processing on data of different modalities, which can convert complex multi-modal data into a feature representation form that is easy to process and analyze, and provides a basis for subsequent multi-modal feature fusion and abnormal behavior classification. From the technical effect, the feature extraction processing can reduce the dimension of the data, improve the processing efficiency, and at the same time retain the key feature information of the data, which helps to improve the accuracy of abnormal behavior recognition.

[0030] In an embodiment, for the first feature extraction process, a pre-trained convolutional neural network model such as ResNet, VGG, etc. can be used. These models have been trained on large-scale image datasets and have strong feature learning capabilities. The visible light image data can be input into the pre-trained convolutional neural network, and the output of the last convolutional layer can be extracted as the first feature vector set. For the second feature extraction process, the infrared image data can be pre-processed first, such as denoising, enhancement, etc., and then a method combining heat gradient feature extraction and heat texture feature extraction can be used to extract representative feature vectors. For the third feature extraction process, an open-source audio processing library such as Librosa can be used, which provides rich audio feature extraction functions. The audio signal data can be input into the Librosa library, and the mel frequency cepstral coefficients and linear predictive cepstral coefficients can be extracted by calling the relevant functions to form the third feature vector set.

[0031] S30: input the first feature vector set, the second feature vector set, and the third feature vector set into a multi-modal feature fusion module, and generate a fusion feature vector set through a weighted fusion algorithm, wherein the weighted fusion algorithm dynamically adjusts the fusion weight according to the confidence weight of each modal feature vector.

[0032] In this application, the multi-modal feature fusion module is a key component for implementing the fusion of different modal feature vectors. Its main function is to integrate the first feature vector set, the second feature vector set, and the third feature vector set to generate a comprehensive fusion feature vector set.

[0033] The weighted fusion algorithm is the core algorithm used by the multi-modal feature fusion module, which can determine the importance of each modality in the fusion process according to the confidence weight of each modal feature vector. The confidence weight reflects the reliability and effectiveness of each modal feature vector in the current environment. For example, in a well-lit environment, the confidence weight of visible light image data may be higher; while in the night or in an environment with more smoke, the confidence weight of infrared image data may be correspondingly increased. Through dynamic adjustment of the fusion weight, the weighted fusion algorithm can fully utilize the advantages of each modal data and improve the quality of the fusion feature vector set.

[0034] The fusion feature vector set is a vector set obtained by weighted fusion of different modal feature vectors, which integrates the feature information of each modal data and can more comprehensively and accurately represent the behavior characteristics in the monitoring scene. In terms of technical effects, the fusion feature vector set can reduce the limitations of single modal data and improve the accuracy and reliability of abnormal behavior recognition. For example, in some complex scenes, single modal data may not be able to accurately identify abnormal behavior, but through multi-modal feature fusion, the complementary nature of different modal data can be utilized to more accurately determine the occurrence of abnormal behavior.

[0035] The embodiment of the present application can fully utilize the advantages of multi-modal data and improve the performance of abnormal behavior recognition by inputting the feature vector sets of different modalities into the multi-modal feature fusion module and generating a fused feature vector set using a weighted fusion algorithm. From the technical effect, the weighted fusion algorithm that dynamically adjusts the fusion weight can adapt to different environmental changes and enhance the robustness and adaptability of the system.

[0036] In an embodiment, a weighted fusion algorithm based on an attention mechanism can be used. The attention mechanism is a method that simulates human attention allocation, which can automatically adjust the fusion weight according to the importance of each modal feature vector. Specifically, the first feature vector set, the second feature vector set and the third feature vector set can be input into an attention network, and the attention network can calculate the attention weight of each modality by learning the feature representation of each modal feature vector. Then, the weighted sum of each modal feature vector is calculated according to the attention weight, and a fused feature vector set is generated. This weighted fusion algorithm based on the attention mechanism can automatically focus on important modal features and improve the quality of the fused feature vector set.

[0037] S40: inputting the fused feature vector set into an abnormal behavior classification model to generate an abnormal behavior classification result, the abnormal behavior classification result including a normal behavior label and an abnormal behavior label.

[0038] In the present application, the abnormal behavior classification model is a trained classifier, and its main task is to determine whether the behavior in the monitoring scene is normal or abnormal according to the input fused feature vector set. The abnormal behavior classification model can use various machine learning or deep learning models, such as support vector machine (SVM), decision tree, convolutional neural network, etc. Support vector machine is a classification model based on statistical learning theory, which separates samples of different classes by finding the optimal classification hyperplane. For example, in a binary classification scenario, the support vector machine can divide the normal behavior samples and abnormal behavior samples into different regions. Decision tree is a classification model based on tree structure, which constructs classification rules by dividing features. For example, in an abnormal behavior recognition scenario, the decision tree can determine whether the behavior is abnormal according to different feature values in the fused feature vector set. Convolutional neural network is a deep learning model that automatically learns feature representation and classification rules through convolutional layers, pooling layers and fully connected layers.

[0039] The abnormal behavior classification result is an output result of the abnormal behavior classification model, and includes a normal behavior label and an abnormal behavior label. The normal behavior label indicates that the behavior in the monitoring scene conforms to the preset normal behavior mode, and the abnormal behavior label indicates that the behavior has an abnormal condition. The technical effect is that whether the behavior in the monitoring scene is abnormal can be quickly and accurately judged, and a basis is provided for subsequent alarm and processing.

[0040] The embodiment of the application can use the classification ability of the model to accurately classify the behavior in the monitoring scene by inputting the fusion feature vector set into the abnormal behavior classification model. In terms of technical effects, the abnormal behavior classification model can improve the efficiency and accuracy of abnormal behavior recognition and reduce the workload of manual judgment.

[0041] In an embodiment, a pre-trained convolutional neural network model can be used as the abnormal behavior classification model. First, the fusion feature vector set is preprocessed, such as normalization processing, to unify the value range of the feature vector to a fixed interval, so as to improve the training effect and stability of the model. Then, the preprocessed fusion feature vector set is input into the pre-trained convolutional neural network. The structure of the convolutional neural network can include multiple convolutional layers, pooling layers and fully connected layers. The convolutional layer is used to extract local feature information of the feature, the pooling layer is used to reduce the dimension of the feature, and the fully connected layer is used to map the extracted feature to the classification result. In the training process, a large amount of labeled data is used to train the model, and the parameters of the model are continuously adjusted so that the model can accurately classify the fusion feature vector set into normal behavior and abnormal behavior. After training, the new fusion feature vector set is input into the trained model, and the model outputs the corresponding abnormal behavior classification result. In terms of technical effects, the pre-trained convolutional neural network model has strong feature learning ability and classification ability, and can accurately identify abnormal behavior in different scenes, thereby improving the intelligent level of the security monitoring system.

[0042] S50: According to the abnormal behavior classification result, an abnormal behavior alarm signal is generated, and the abnormal behavior alarm signal is sent to the security monitoring terminal to trigger an alarm operation.

[0043] In this application, the abnormal behavior alarm signal is a warning information generated according to the abnormal behavior classification result, which is used to inform the security monitoring personnel that there is an abnormal behavior in the monitoring scene. When the abnormal behavior classification result is an abnormal behavior label, the system will automatically generate an abnormal behavior alarm signal. The abnormal behavior alarm signal can contain various forms of information, such as text description, image information, sound prompt, etc. For example, the text description can detail the type, time and location of the abnormal behavior; the image information can be related visible light image or infrared image, which can intuitively show the specific situation of the abnormal behavior; the sound prompt can issue an alarm sound through the loudspeaker to attract the attention of the monitoring personnel.

[0044] The security monitoring terminal is a device that receives and processes the abnormal behavior alarm signal, which is usually a special monitoring platform or software system. After receiving the abnormal behavior alarm signal, the security monitoring terminal will trigger the corresponding alarm operation. The alarm operation can include popping up an alarm prompt window on the monitoring screen, issuing a sound and light alarm, recording alarm information, etc. For example, when the alarm prompt window pops up on the monitoring screen, it will display the detailed information and related images of the abnormal behavior, so that the monitoring personnel can timely understand the situation of the abnormal behavior. At the same time, the sound and light alarm will attract the attention of the monitoring personnel, reminding them to take timely measures. Recording alarm information can save the related data of the abnormal behavior for subsequent analysis and processing.

[0045] The embodiments of the present application can timely and effectively convey the abnormal behavior information to the monitoring personnel by generating the abnormal behavior alarm signal according to the abnormal behavior classification result and sending it to the security monitoring terminal, so that they can respond quickly. From the technical effect, the timely alarm operation can reduce the loss caused by the abnormal behavior and improve the security and reliability of the security monitoring system.

[0046] In an embodiment, in order to ensure that the abnormal behavior alarm signal can be accurately and timely sent to the security monitoring terminal, multiple communication methods can be used for data transmission. For example, wireless network communication technology such as Wi-Fi, 4G or 5G network can be used to send the abnormal behavior alarm signal to the security monitoring terminal through the network. At the same time, multiple backup communication links can be set to prevent the failure of a single communication link from causing the alarm signal to be unable to be transmitted. In terms of security monitoring terminals, special monitoring software can be developed, which has the function of receiving, processing and displaying abnormal behavior alarm signals in real time. The software can classify and manage alarm signals, and provide different levels of alarm prompts according to the severity of abnormal behavior. For example, for serious abnormal behavior such as violent conflict, theft, etc., high-intensity sound and light alarms and emergency prompts can be used; for minor abnormal behavior such as personnel entering in violation of rules, etc., relatively mild prompt methods can be used. In addition, the monitoring software can also record the detailed content of the alarm information, including the time, location, type and handling of the abnormal behavior, etc., for subsequent inquiry and analysis.

[0047] In an embodiment, referring to Figure 3 , step S30 can include steps S301-S305, which are described in detail as follows: Step S301: calculating a first confidence weight value of the first feature vector set, a second confidence weight value of the second feature vector set, and a third confidence weight value of the third feature vector set.

[0048] In this application, the confidence weight value reflects the reliability and importance of each feature vector set in the multi-modal feature fusion process. The first confidence weight value is used to measure the credibility of the first feature vector set, the second confidence weight value is used to measure the credibility of the second feature vector set, and the third confidence weight value is used to measure the credibility of the third feature vector set. By calculating these confidence weight values, the fusion weight can be reasonably allocated according to the actual situation of each modal feature vector, and the quality of the fused feature vector set can be improved.

[0049] In an embodiment, referring to Figure 4 , step S301 can include steps S3011-S3013, which are described in detail as follows: Step S3011: obtaining a first representation intensity value of each feature vector in the first feature vector set, a second representation intensity value of each feature vector in the second feature vector set, and a third representation intensity value of each feature vector in the third feature vector set.

[0050] In the present application, the characteristic intensity value is an index for measuring the information intensity represented by the feature vector. The first characteristic intensity value reflects the richness and significance of the information carried by each feature vector in the first feature vector set. For example, in the feature vectors of visible light image data, some feature vectors may represent key targets in the image, and their first characteristic intensity values may be higher; while some feature vectors may only represent background information, and their first characteristic intensity values may be lower. Similarly, the second characteristic intensity value and the third characteristic intensity value reflect the information intensity of each feature vector in the second feature vector set and the third feature vector set, respectively. Obtaining the characteristic intensity value of each feature vector helps to evaluate the overall intensity of each feature vector set subsequently.

[0051] As a possible implementation, for a feature vector , its norm can be used as the characteristic intensity value, and the calculation formula is: wherein is the i-th element of the feature vector , and n is the dimension of the feature vector.

[0052] Step S3012: According to the first characteristic intensity value, the second characteristic intensity value and the third characteristic intensity value, respectively calculate the first characteristic intensity mean value of the first feature vector set, the second characteristic intensity mean value of the second feature vector set and the third characteristic intensity mean value of the third feature vector set.

[0053] In the present application, the characteristic intensity mean value is an average measure of the characteristic intensity values of all feature vectors in a feature vector set. The first characteristic intensity mean value reflects the overall information intensity level of the first feature vector set. By calculating the first characteristic intensity mean value, the conditions of all feature vectors in the first feature vector set can be considered comprehensively to obtain an index that can represent the overall characteristics of the set. Similarly, the second characteristic intensity mean value and the third characteristic intensity mean value reflect the overall information intensity level of the second feature vector set and the third feature vector set, respectively. Calculating the characteristic intensity mean value can provide basic data for subsequent calculation of the confidence weight value.

[0054] In an embodiment, the calculation formula of the first characteristic intensity mean value is: , wherein is the first characteristic intensity value of the j-th feature vector in the first feature vector set, and m1 is the number of feature vectors in the first feature vector set. Similarly, the second characteristic intensity mean value and the third characteristic intensity mean value are: , wherein m2 and m3 are the number of eigenvectors in the second eigenvector set and the third eigenvector set respectively.

[0055] Step S3013: calculating the first confidence weight value, the second confidence weight value and the third confidence weight value respectively according to the first representation intensity mean value, the second representation intensity mean value and the third representation intensity mean value.

[0056] In the present application, the purpose of calculating the confidence weight value according to the representation intensity mean value is to enable each eigenvector set to be reasonably allocated weight according to its actual information intensity in multi-modal feature fusion. Generally speaking, the eigenvector set with a higher representation intensity mean value will also have a relatively higher confidence weight value, because it carries more reliable and important information.

[0057] In an embodiment, step S3013 can include steps A-E, which are described in detail as follows: A: calculating the deviation value between the first representation intensity mean value and a preset first reference value, the deviation value between the second representation intensity mean value and a preset second reference value, and the deviation value between the third representation intensity mean value and a preset third reference value.

[0058] In the present application, the preset first reference value, the preset second reference value and the preset third reference value are reference values preset in advance, which respectively represent the representation intensity mean value of the first eigenvector set, the second eigenvector set and the third eigenvector set in an ideal case. By calculating the deviation value between the representation intensity mean value and the reference value, the difference between the actual information intensity of each eigenvector set and the ideal case can be understood. The smaller the deviation value, the closer the information intensity of the eigenvector set to the ideal level, and the relatively higher the reliability.

[0059] In an embodiment, the deviation value between the first representation intensity mean value and the preset first reference value is . The calculation formula is . Similarly, the deviation value between the second representation intensity mean value and the preset second reference value is , and the deviation value between the third representation intensity mean value and the preset third reference value is . .

[0060] B: determining the preliminary value range of the first confidence weight value according to the deviation value between the first representation intensity mean value and the first reference value.

[0061] ​In the present application, the deviation value between the first characteristic intensity mean value and the first reference value reflects the deviation degree of the actual information intensity of the first feature vector set from the ideal level. According to the deviation value, the value range of the first confidence weight value can be preliminarily determined. Generally speaking, the smaller the deviation value is, the higher the preliminary value range of the first confidence weight value will be, because the information of the feature vector set is more reliable. For example, if the deviation value is very small, it means that the information intensity of the first feature vector set is close to the ideal level, so the preliminary value range of the first confidence weight value can be in a higher interval.

[0062] C: determining a preliminary value range of the second confidence weight value according to the deviation value between the second characteristic intensity mean value and the second reference value.

[0063] Similarly, the deviation value between the second characteristic intensity mean value and the second reference value determines the preliminary value range of the second confidence weight value. The smaller the deviation value is, the higher the preliminary value range of the second confidence weight value will be, indicating that the information of the second feature vector set is more reliable.

[0064] D: determining a preliminary value range of the third confidence weight value according to the deviation value between the third characteristic intensity mean value and the third reference value.

[0065] The deviation value between the third characteristic intensity mean value and the third reference value is used to determine the preliminary value range of the third confidence weight value. The smaller the deviation value is, the higher the preliminary value range of the third confidence weight value will be, indicating that the information of the third feature vector set is more reliable.

[0066] E: normalizing the preliminary value range of the first confidence weight value, the preliminary value range of the second confidence weight value and the preliminary value range of the third confidence weight value to obtain the first confidence weight value, the second confidence weight value and the third confidence weight value.

[0067] In the present application, the normalization process is the process of converting the preliminary value range to a specific confidence weight value. The purpose of the normalization process is to make the sum of the first confidence weight value, the second confidence weight value and the third confidence weight value equal to 1, so that the weight distribution of each feature vector set in the multi-modal feature fusion process can be reasonable. Through the normalization process, the final first confidence weight value, the second confidence weight value and the third confidence weight value can be obtained, which are used for subsequent weighted fusion calculation.

[0068] Step 302: determining an initial fusion weight distribution table of the multi-modal feature fusion module according to the first confidence weight value, the second confidence weight value and the third confidence weight value.

[0069] In the present application, the initial fusion weight distribution table is formulated according to the first confidence weight value, the second confidence weight value and the third confidence weight value, and is used to define the initial weight distribution of each feature vector set in the multi-modal feature fusion process. The initial fusion weight distribution table can be presented in the form of a table, in which the weight values corresponding to the first feature vector set, the second feature vector set and the third feature vector set are explicitly listed. By determining the initial fusion weight distribution table, a basis can be provided for subsequent weighted calculation.

[0070] Step S303: For each feature vector in the first feature vector set, a weighted value of the feature vector is calculated according to the first confidence weight value in the initial fusion weight distribution table.

[0071] In the present application, the weighted value is the result of weighted processing of the feature vector. For each feature vector in the first feature vector set, weighted calculation is performed according to the first confidence weight value in the initial fusion weight distribution table. Specifically, the element value of each feature vector is multiplied by the first confidence weight value to obtain the weighted value of the feature vector. By calculating the weighted value, the role of important features in the first feature vector set can be highlighted, while the influence of unimportant features can be reduced.

[0072] Step S304: For each feature vector in the second feature vector set and the third feature vector set, a corresponding weighted value is calculated according to the second confidence weight value and the third confidence weight value in the initial fusion weight distribution table, respectively.

[0073] Similarly, for each feature vector in the second feature vector set and the third feature vector set, weighted calculation is performed according to the second confidence weight value and the third confidence weight value in the initial fusion weight distribution table, respectively. The element value of each feature vector is multiplied by the corresponding confidence weight value to obtain the corresponding weighted value. In this way, the weights can be reasonably distributed according to the credibility of each feature vector set, and the quality of the fused feature vector set can be improved.

[0074] Step S305: All weighted values are combined according to a preset rule to generate the fused feature vector set.

[0075] In the present application, the preset rule is a pre-set combination mode for merging all the weighted values into a fusion feature vector set. The preset rule can be designed according to specific application scenarios and requirements. For example, a simple vector splicing method can be used to splice the weighted values of the first feature vector set, the second feature vector set and the third feature vector set in turn to form a new fusion feature vector set. A weighted summation method can also be used to weight and sum the weighted values of each feature vector set to obtain a comprehensive fusion feature vector set. By combining the weighted values according to the preset rule, a fusion feature vector set that can integrate the feature information of each modality can be obtained.

[0076] In an embodiment, for obtaining the characteristic intensity value in step S301, the length of the feature vector can be used as the characteristic intensity value. For each feature vector, the square root of the sum of squares of its elements is calculated to obtain the length of the feature vector as its characteristic intensity value. For determining the initial fusion weight distribution table in step S302, the first confidence weight value, the second confidence weight value and the third confidence weight value can be directly filled into the corresponding positions in the table. For calculating the weighted values in steps S303-S304, a matrix operation method can be used to multiply the feature vector matrix with the corresponding confidence weight values. For combining the weighted values in step S305, a vector splicing method can be used to splice the weighted values of each feature vector set together using an array splicing function in a programming language to generate a fusion feature vector set.

[0077] In an embodiment, with reference to Figure 5 The method provided in the present application further includes, before step S50: Step S61: Obtain a training data set, wherein the training data set includes first sample data labeled as normal behavior and second sample data labeled as abnormal behavior.

[0078] In the present application, the training data set is the basic data for training the abnormal behavior classification model. The first sample data is labeled as normal behavior, which contains multi-modal data collected under normal circumstances, such as visible light images, infrared images and audio signals of normal personnel, etc. The second sample data is labeled as abnormal behavior, which contains multi-modal data under various abnormal behavior conditions, such as visible light images, infrared images and audio signals under violent conflict, theft and other scenes. By obtaining a training data set containing normal behavior and abnormal behavior samples, the abnormal behavior classification model can learn the feature differences between normal behavior and abnormal behavior, thereby having the ability to accurately classify.

[0079] Step S62: input the first sample data and the second sample data into the initial abnormal behavior classification model for training to generate a trained abnormal behavior classification model.

[0080] In this application, the initial abnormal behavior classification model is an untrained classification model, which can adopt model structures such as support vector machines, decision trees, and convolutional neural networks as described above. The first sample data and the second sample data are input into the initial abnormal behavior classification model, and the model will learn the feature patterns of normal behavior and abnormal behavior according to the input data. During the training process, the model will continuously adjust its parameters to continuously improve the classification accuracy of the training data. After a certain number of training iterations, a trained abnormal behavior classification model is generated, which can accurately classify new fusion feature vector sets.

[0081] Step S63: In the training process, the cross-validation method is used to evaluate the classification performance of the initial abnormal behavior classification model, and the hyperparameter configuration of the initial abnormal behavior classification model is adjusted according to the evaluation results.

[0082] In this application, the cross-validation method is a technique for evaluating model performance. It divides the training data set into multiple subsets, then uses one subset as the validation set and the remaining subsets as the training set for training and validation. By repeating this process multiple times, the classification performance indicators of the model on different subsets can be obtained, such as accuracy, recall, F1 value, etc. According to these performance indicators, the classification performance of the initial abnormal behavior classification model can be evaluated. If the evaluation result is not ideal, it means that the hyperparameter configuration of the model may not be appropriate. Hyperparameters are parameters that need to be set in advance before model training, such as learning rate, number of iterations, regularization parameter, etc. Adjusting the hyperparameter configuration of the initial abnormal behavior classification model according to the evaluation results can improve the classification performance and generalization ability of the model.

[0083] In an embodiment, for step S61 to obtain the training data set, multi-modal data can be collected from multiple security monitoring scenes and labeled by professionals to distinguish normal behavior and abnormal behavior samples. For step S62 to train the model, a deep learning framework such as TensorFlow or PyTorch can be used to input the first sample data and the second sample data into the initial abnormal behavior classification model for training. For step S63 cross-validation, the cross-validation function in the sklearn library can be used to divide the training data set into multiple subsets for cross-validation, and the hyperparameters of the model can be adjusted according to the validation results.

[0084] Suppose the data set is divided into k folds by cross-validation, and the accuracy of the i-th fold validation set is Then the average accuracy is : The hyperparameters of the initial abnormal behavior classification model can be adjusted according to indicators such as average accuracy For example, an update formula of the gradient descent type is used: wherein is a learning rate, is a new parameter, denotes a parameter of the last iteration, is a loss function, is a gradient with respect to the hyperparameter .

[0085] In an embodiment, with reference to Figure 6 , the application further includes a step of pre-processing the first sample data and the second sample data in the process of constructing the training data set, to ensure the quality of the data input into the initial abnormal behavior classification model, as follows: Step S71: Obtain the original annotation information of each sample in the first sample data and the second sample data, and verify the accuracy of the original annotation information.

[0086] In the present application, the original annotation information is the basis for classifying and annotating sample data, and it contains label information indicating whether the sample is normal behavior or abnormal behavior. It is very important to verify the accuracy of the original annotation information, because incorrect annotation information will cause the model to learn incorrect feature patterns, thereby affecting the classification performance of the model. During the verification process, the accuracy of the original annotation information can be ensured through manual inspection, data comparison, etc. For example, for some samples with ambiguous or questionable annotations, multiple professionals can re-annotate and confirm them.

[0087] For sample visible light image data , the normalization formula can be expressed as: wherein is the minimum value in the image data, is the maximum value.

[0088] Similarly, for sample infrared image data and sample audio signal data, similar normalization formulas are used.

[0089] Step S72: For each sample, sample multi-modal data corresponding to the sample is extracted according to the original annotation information, and the multi-modal data includes sample visible light image data, sample infrared image data, and sample audio signal data.

[0090] In the present application, the corresponding sample multi-modal data is extracted according to the original annotation information, so as to ensure that the data of each sample is consistent with the annotation information. For each sample, the corresponding visible light image data, infrared image data and audio signal data are extracted from the original data. For example, if the original annotation information of a sample is abnormal behavior, then the visible light image data, infrared image data and audio signal data corresponding to the sample in the abnormal behavior scene are extracted. By accurately extracting the sample multi-modal data, accurate data basis can be provided for subsequent normalization processing and model training.

[0091] Step S73: The sample visible light image data, sample infrared image data and sample audio signal data are respectively normalized to generate standardized first sample data and second sample data.

[0092] In the present application, normalization processing is a process of converting data to a unified scale and range. Normalizing the sample visible light image data, sample infrared image data and sample audio signal data respectively can eliminate the dimensional differences between different data, so that the data is comparable. For the sample visible light image data, the pixel value of the image can be normalized to the interval [0, 1]; for the sample infrared image data, the thermal radiation value thereof can be normalized; and for the sample audio signal data, the audio amplitude thereof can be normalized. Through normalization processing, standardized first sample data and second sample data are generated, which helps to improve the training effect and stability of the model and avoids the problem of difficult model training or slow convergence speed caused by too large data scale difference.

[0093] Step S74: The standardized first sample data and second sample data are divided into a training data set, a validation data set and a test data set, and the distribution proportion of each type of sample in each data set is ensured to be balanced during the division process.

[0094] In this application, the standardized sample data is divided into training data set, validation data set and test data set, in order to comprehensively evaluate and optimize the model. The training data set is used for the training of the model, so that the model can learn the feature mode of the sample; the validation data set is used to evaluate the performance of the model during the training process, and adjust the hyperparameters of the model; the test data set is used to evaluate the final performance of the model after the training of the model is completed. In the division process, it is very important to ensure that the distribution proportion of each type of sample (normal behavior sample and abnormal behavior sample) in each data set is balanced. If the proportion of a certain type of sample in a certain data set is too high or too low, it will lead to insufficient learning of the model for this type of sample or overfitting. For example, if the proportion of abnormal behavior samples in the training data set is too low, the model may not be able to fully learn the characteristics of abnormal behavior, thereby affecting its ability to identify abnormal behavior. By balancing the distribution proportion of each type of sample in each data set, the generalization ability and classification accuracy of the model can be improved.

[0095] Suppose the number of first sample data after standardization is , the number of second sample data is , the proportion of training set is , the proportion of validation set is , and the proportion of test set is , then: The number of first sample data in the training set is , and the number of second sample data is The number of first sample data in the validation set is , and the number of second sample data is The number of first sample data in the test set is , and the number of second sample data is Step S75: Based on the validation data set and test data set, the performance of the initial abnormal behavior classification model is preliminarily evaluated, and the sample weight configuration of the training data set is adjusted according to the evaluation result.

[0096] ​In this application, the performance of the initial abnormal behavior classification model is preliminarily evaluated by using the verification data set and the test data set, so as to understand the performance of the model on different data sets. The evaluation indicators can include accuracy, recall rate, F1 value, etc. If the evaluation result shows that the performance of the model on some categories is poor, it means that the model does not learn enough samples of these categories. At this time, the sample weight configuration of the training data set can be adjusted according to the evaluation result. For example, if the identification accuracy of the model for abnormal behavior samples is low, the weight of the abnormal behavior samples in the training data set can be increased, so that the model pays more attention to the features of these samples, thereby improving the identification ability of the abnormal behavior. By adjusting the sample weight configuration, the training process of the model can be optimized, and the performance of the model can be improved.

[0097] Let the loss on the verification set be , and the loss on the test set be The weight of the jth sample in the training data set can be adjusted according to the loss value , for example: wherein, , , is a preset coefficient, is the initial weight of the jth sample.

[0098] In an embodiment, for step S71 of verifying the accuracy of the original annotation information, an annotation audit team can be established, and professional personnel can be used to audit and correct the annotation information. For step S72 of extracting sample multi-modal data, a data management system can be used to filter the corresponding sample multi-modal data from the database according to the original annotation information. For step S73 of normalization processing, for the sample visible light image data, an image normalization function can be used to normalize the pixel value by dividing by 255; for the sample infrared image data, a statistical method can be used to calculate the mean and standard deviation, and then the standardization processing is performed; for the sample audio signal data, a normalization function in the audio processing library can be used for processing. For step S74 of dividing the data set, a random sampling method can be used to divide the standardized sample data into a training data set, a verification data set and a test data set according to a certain proportion, and a stratified sampling method is used to ensure that the distribution proportion of each type of sample is balanced. For step S75 of adjusting the sample weight configuration, the evaluation results of the verification data set and the test data set can be used to dynamically adjust the weights of each type of sample in the training data set by using a weight adjustment algorithm.

[0099] In an embodiment, the weighting fusion algorithm of the multi-modal feature fusion module described in the application further includes a mechanism for dynamically adjusting the confidence weight value to adapt to environmental changes in different monitoring periods, as described in Figure 7The method of the present application further comprises: Step S81: In the target monitoring period, according to the real-time acquisition of the multi-modal data, the environmental interference factor value in each monitoring period is calculated.

[0100] In the present application, the target monitoring period is a time period set in advance, and in this time period, the security camera continuously acquires multi-modal data. The environmental interference factor value is used to measure the degree of interference of the environment on the acquisition of multi-modal data in each monitoring period. Different environmental factors will have different degrees of influence on visible light image data, infrared image data and audio signal data. For example, too strong or too weak light intensity will affect the clarity of the visible light image; too high or too low temperature will affect the thermal radiation characteristics of the infrared image; too large noise intensity will interfere with the acquisition of the audio signal. By calculating the environmental interference factor value, the degree of interference of the environment on the multi-modal data can be quantified, which provides a basis for subsequent determination of the credibility level of the data and adjustment of the confidence weight value.

[0101] In an embodiment, referring to Figure 8 Step S81 can include steps S810-S812, which are described in detail below: Step S810: Obtain the light intensity value, temperature value and noise intensity value of the environment in which the security camera is located in the target monitoring period.

[0102] In the present application, the light intensity value reflects the brightness of the light in the environment, which can be measured by a light sensor. The temperature value is a thermal state indicator of the environment, which is obtained by a temperature sensor. The noise intensity value measures the degree of noise in the environment, which is acquired by a noise sensor. Obtaining these environmental parameter values is the basis for calculating the environmental interference factor value. For example, in an outdoor security monitoring scene, the light intensity value will have a big difference between day and night, the temperature value will change in different seasons and time periods, and the noise intensity value will be different in different time periods and locations. Accurate acquisition of these environmental parameter values can more accurately evaluate the influence of the environment on the acquisition of multi-modal data.

[0103] Step S811: According to the light intensity value, temperature value and noise intensity value, the comprehensive environmental parameter in each monitoring period is calculated.

[0104] In the present application, the comprehensive environmental parameter is an index obtained by comprehensively considering the illumination intensity value, the temperature value and the noise intensity value. It can more comprehensively reflect the influence of the environment on the multi-modal data acquisition. When calculating the comprehensive environmental parameter, the weights of the environmental parameters need to be considered. Different environmental parameters may have different influences on different modal data. For example, the illumination intensity value has a greater influence on the visible light image data, and the temperature value has a more critical influence on the infrared image data. By reasonably setting the weights of the environmental parameters, a more accurate comprehensive environmental parameter can be calculated.

[0105] In an embodiment, referring to Figure 9 , step S811 can include steps S8110-S8113, which are described in detail as follows: Step S8110: Obtain the illumination intensity value sequence, the temperature value sequence and the noise intensity value sequence of each monitoring period in the target monitoring period.

[0106] In the present application, the illumination intensity value sequence is an ordered set of illumination intensity values recorded in each monitoring period in the target monitoring period, and the temperature value sequence and the noise intensity value sequence are the same. Obtaining these sequence data can more comprehensively understand the changes of the environmental parameters in different monitoring periods. For example, by analyzing the illumination intensity value sequence, it can be found that the change rule of the illumination intensity in a day and whether there is an abnormal illumination fluctuation. These sequence data provide a data basis for subsequent smoothing filtering processing and mean value calculation.

[0107] Step S8111: Perform smoothing filtering processing on the illumination intensity value sequence, the temperature value sequence and the noise intensity value sequence respectively to eliminate the influence of abnormal values.

[0108] In the present application, the smoothing filtering processing is a data processing method for removing abnormal values and noise in the data sequence. In the environmental parameter acquisition process, some abnormal values may be generated due to sensor failure, external interference and other reasons. These abnormal values will affect the calculation accuracy of the subsequent comprehensive environmental parameter. Through the smoothing filtering processing, the data sequence can be made more smooth, and the influence of abnormal values can be reduced. For example, the moving average filtering method can be used to calculate the average value of a certain number of data points before and after each data point in the illumination intensity value sequence as the smoothed value of the data point.

[0109] In an embodiment, simple moving average filtering can be adopted. For the illumination intensity value sequence , the smoothed value is: Wherein M is the window size, and i≥M.

[0110] Similarly, the temperature value sequence and the noise intensity value sequence are processed in a similar manner.

[0111] Step S8112: According to the smoothed sequence of illumination intensity values, the sequence of temperature values and the sequence of noise intensity values, the average illumination intensity, the average temperature and the average noise intensity in each monitoring period are calculated.

[0112] In this application, the average illumination intensity, the average temperature and the average noise intensity are the average values of the smoothed sequence of illumination intensity values, the sequence of temperature values and the sequence of noise intensity values in each monitoring period. Calculating these average values can more stably reflect the environmental parameter level in each monitoring period. For example, the average illumination intensity can represent the average illumination intensity in the monitoring period, and the average temperature can represent the average temperature in the monitoring period. These average value data provide a more reliable basis for calculating the comprehensive environmental parameter.

[0113] Step S8113: Based on the average illumination intensity, the average temperature and the average noise intensity, combined with the preset weight coefficient, the comprehensive environmental parameter in each monitoring period is calculated.

[0114] In this application, the preset weight coefficient is preset according to the influence degree of each environmental parameter on multi-modal data acquisition. For example, assuming that the influence of illumination intensity on visible light image data is larger, its weight coefficient can be set higher; and the influence of noise intensity on audio signal data is larger, its weight coefficient can also be increased accordingly. Multiply the average illumination intensity, the average temperature and the average noise intensity by the corresponding weight coefficient respectively, and then add the results, which can obtain the comprehensive environmental parameter in each monitoring period. The comprehensive environmental parameter can more comprehensively reflect the influence of the environmental condition in the monitoring period on multi-modal data acquisition.

[0115] Let the preset illumination intensity weight coefficient be , the temperature weight coefficient be , the noise intensity weight coefficient be , and . The calculation formula of the comprehensive environmental parameter Step S812: Based on the comprehensive environmental parameter, combined with the pre-set environmental interference threshold range, the environmental interference factor value in each monitoring period is determined.

[0116] ​In the present application, the pre-set environmental interference threshold range is determined according to actual experience and experimental data, and is used to judge the degree of environmental interference. The comprehensive environmental parameter is compared with the environmental interference threshold range. If the comprehensive environmental parameter exceeds the threshold range, it indicates that the degree of environmental interference is large, and the corresponding environmental interference factor value will be high. If the comprehensive environmental parameter is within the threshold range, it indicates that the degree of environmental interference is small, and the environmental interference factor value will be low. By determining the environmental interference factor value, the degree of environmental interference in each monitoring period can be quantified.

[0117] The pre-set environmental interference threshold range is When , the environmental interference factor value can be expressed as: When , the environmental interference factor value is: When , the environmental interference factor value is: Step S82: Based on the environmental interference factor value, determine the credibility level of the visible light image data, infrared image data and audio signal data in each monitoring period.

[0118] In the present application, the credibility level is used to evaluate the reliability of the visible light image data, infrared image data and audio signal data in each monitoring period. The greater the environmental interference factor value, the greater the environmental interference to data acquisition, and the lower the corresponding data credibility level. For example, when the environmental interference factor value is high, the visible light image may be blurred due to insufficient light, and its credibility level will be reduced; the infrared image may be affected by temperature fluctuations, resulting in inaccurate thermal radiation characteristics, and its credibility level will also decrease; the audio signal may be covered by noise, and the credibility level will also decrease. According to the environmental interference factor value, the credibility level of the data is determined, which can provide a basis for subsequent adjustment of the confidence weight value.

[0119] The mapping relationship between the environmental interference factor value and the credibility level can be established. Assuming that the credibility level is divided into high, medium and low three levels, respectively represented by . Set the threshold values and , then: When , the credibility level ; when When the credibility level ; when When the credibility level .

[0120] Step S83: dynamically adjusting the confidence weight values ​​of the first feature vector set, the second feature vector set, and the third feature vector set according to the credibility level.

[0121] In this application, the confidence weight value is dynamically adjusted according to the credibility level in order to make the multimodal feature fusion process more reasonable. When the credibility level of a certain modal data is low, the confidence weight value of the corresponding feature vector set will be reduced to reduce the impact of the modal data on the fusion result; when the credibility level of a certain modal data is high, the confidence weight value of its feature vector set will be increased to give full play to the advantages of the modal data. For example, if the credibility level of the visible light image data is low during a certain monitoring period, the confidence weight value of the first feature vector set will be reduced accordingly; and the credibility level of the infrared image data is high, the confidence weight value of the second feature vector set will be increased. By dynamically adjusting the confidence weight value, the fused feature vector set can more accurately reflect the actual situation of the monitoring scene and improve the accuracy of abnormal behavior recognition.

[0122] In one embodiment, for obtaining environmental parameter values ​​in step S810, light sensors, temperature sensors, and noise sensors can be installed near the security camera to collect these environmental parameter values ​​in real time. For calculating the comprehensive environmental parameters in step S811, while acquiring sequence data in step S8110, the data acquisition system can periodically record the environmental parameter values. For smoothing filtering in step S8111, a filtering function from the Python scipy library can be used. For calculating the mean in step S8112, a mean calculation function from the numpy library can be used. For calculating the comprehensive environmental parameters in step S8113, a weighted summation is performed based on preset weight coefficients. For determining the environmental interference factor value in step S812, the comprehensive environmental parameters can be compared with a pre-set environmental interference threshold range, and the environmental interference factor value can be determined based on the comparison result. For determining the credibility level in step S82, a credibility level mapping table can be established to search for the corresponding credibility level based on the environmental interference factor value. Regarding step S83 of adjusting the confidence weight values, the confidence weight values ​​of the first feature vector set, the second feature vector set, and the third feature vector set may be dynamically adjusted according to the correspondence between the credibility levels and the confidence weight values.

[0123] The security camera abnormal behavior recognition method based on multi-modal fusion of the present application has significant and multi-dimensional technical effects. From the data acquisition and processing level, by acquiring multi-modal data such as visible light image data, infrared image data and audio signal data, the behavior characteristics in the monitoring scene can be fully and stereoscopically captured. The feature extraction is respectively performed on each modal data, and a dynamic weighted fusion algorithm is used, fully considering the reliability difference of different modal data in different environments, greatly improving the accuracy and effectiveness of the fusion feature vector set. At the same time, in the training process of the abnormal behavior classification model, through strict data preprocessing, cross-validation and hyperparameter adjustment operations, the classification performance and generalization ability of the model are effectively improved, ensuring that the abnormal behavior can also be accurately recognized in complex and variable actual scenes.

[0124] From the actual application and system performance level, the dynamic adjustment mechanism of the method gives the system strong environmental adaptability. The weighted fusion algorithm of the multi-modal feature fusion module can dynamically adjust the confidence weight value according to the environmental interference factor value, so that the system can flexibly respond to environmental changes in different monitoring periods, and reduce the influence of environmental interference on the recognition result. Finally, according to the accurate abnormal behavior classification result, an alarm signal is generated in time and sent to the security monitoring terminal, which can discover and warn abnormal behavior at the first time, provides strong technical support for protecting public safety and property safety, and significantly improves the overall efficiency and reliability of the security monitoring system.

[0125] Correspondingly, in order to better implement the above method, the embodiment of the present application also provides a security camera abnormal behavior recognition system 90 based on multi-modal fusion, wherein, as shown in Figure 10 the security camera abnormal behavior recognition system 90 based on multi-modal fusion includes: An acquisition module 901 is configured to acquire multi-modal data collected by a security camera in a target monitoring period, wherein the multi-modal data includes visible light image data, infrared image data and audio signal data. A feature extraction module 902 is configured to perform first feature extraction processing on the visible light image data to generate a first feature vector set, perform second feature extraction processing on the infrared image data to generate a second feature vector set, and perform third feature extraction processing on the audio signal data to generate a third feature vector set. A fusion module 903 is configured to input the first feature vector set, the second feature vector set and the third feature vector set into a multi-modal feature fusion module, and generate a fusion feature vector set by a weighted fusion algorithm, wherein the weighted fusion algorithm dynamically adjusts the fusion weight according to the confidence weight of each modal feature vector. The anomaly classification module 904 is configured to input the set of fusion feature vectors into an abnormal behavior classification model to generate an abnormal behavior classification result, the abnormal behavior classification result including a normal behavior label and an abnormal behavior label. The alarm module 905 is configured to generate an abnormal behavior alarm signal according to the abnormal behavior classification result, and send the abnormal behavior alarm signal to a security monitoring terminal to trigger an alarm operation. The implementation of each of the above modules can be specifically referred to the foregoing method embodiments, which will not be repeated here. The technical effects of the modules and the device are described with reference to the foregoing method embodiments.

[0126] It should be noted that, in specific implementation, each of the above modules can be combined and integrated in one or more modules, or can be implemented as an independent entity. In addition, the above modules can be implemented in the form of hardware or software function modules. The integrated modules, if implemented in the form of software function modules and sold or used as independent products, can also be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0127] In several embodiments provided in the present application, it should be understood that the disclosed system and method can be implemented in other ways. For example, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0128] The above only describes the preferred embodiments of the present application, and of course cannot limit the scope of the rights of the present application, therefore, the equivalent changes made according to the claims of the present application still fall within the scope of the present application.

Claims

1. A method for identifying abnormal behavior of security cameras based on multimodal fusion, characterized in that: The following steps are involved: Acquire multimodal data collected by the security camera during a target monitoring period, the multimodal data including visible light image data, infrared image data, and audio signal data; performing a first feature extraction process on the visible light image data to generate a first feature vector set, performing a second feature extraction process on the infrared image data to generate a second feature vector set, and performing a third feature extraction process on the audio signal data to generate a third feature vector set; Inputting the first feature vector set, the second feature vector set, and the third feature vector set into a multimodal feature fusion module, and generating a fused feature vector set using a weighted fusion algorithm, wherein the weighted fusion algorithm dynamically adjusts the fusion weight according to the confidence weight of each modal feature vector; Inputting the fused feature vector set into an abnormal behavior classification model to generate an abnormal behavior classification result, wherein the abnormal behavior classification result includes a normal behavior mark and an abnormal behavior mark; According to the abnormal behavior classification result, an abnormal behavior alarm signal is generated, and the abnormal behavior alarm signal is sent to the security monitoring terminal to trigger an alarm operation.

2. The method according to claim 1, characterized in that Inputting the first feature vector set, the second feature vector set, and the third feature vector set into a multimodal feature fusion module, and generating a fused feature vector set through a weighted fusion algorithm, including: Calculating a first confidence weight value of the first feature vector set, a second confidence weight value of the second feature vector set, and a third confidence weight value of the third feature vector set; Determining an initial fusion weight distribution table of the multimodal feature fusion module according to the first confidence weight value, the second confidence weight value, and the third confidence weight value; For each feature vector in the first feature vector set, calculating a weighted value of the feature vector according to the first confidence weight value in the initial fusion weight allocation table; For each eigenvector in the second eigenvector set and the third eigenvector set, calculating a corresponding weighted value according to the second confidence weight value and the third confidence weight value in the initial fusion weight allocation table; All weighted values ​​are combined according to preset rules to generate the fused feature vector set.

3. The method according to claim 2, characterized in that Calculating a first confidence weight value of the first feature vector set, a second confidence weight value of the second feature vector set, and a third confidence weight value of the third feature vector set, including: Obtaining a first characterization intensity value of each feature vector in the first feature vector set, a second characterization intensity value of each feature vector in the second feature vector set, and a third characterization intensity value of each feature vector in the third feature vector set; Calculating, according to the first characterization intensity value, the second characterization intensity value, and the third characterization intensity value, a first characterization intensity mean of the first feature vector set, a second characterization intensity mean of the second feature vector set, and a third characterization intensity mean of the third feature vector set, respectively; The first confidence weight value, the second confidence weight value, and the third confidence weight value are calculated respectively according to the first characterization strength mean, the second characterization strength mean, and the third characterization strength mean.

4. The method according to claim 3, characterized in that Calculating the first confidence weight value, the second confidence weight value, and the third confidence weight value respectively according to the first characterization strength mean, the second characterization strength mean, and the third characterization strength mean, including: Calculating a deviation between the first characterization intensity mean and a preset first reference value, a deviation between the second characterization intensity mean and a preset second reference value, and a deviation between the third characterization intensity mean and a preset third reference value; Determining a preliminary value range of the first confidence weight value according to a deviation between the first characterization strength mean and a first reference value; Determining a preliminary value range of the second confidence weight value according to a deviation between the second characterization strength mean and a second reference value; Determining a preliminary value range of the third confidence weight value according to a deviation value between the third characterization strength mean and a third reference value; The preliminary value range of the first confidence weight value, the preliminary value range of the second confidence weight value, and the preliminary value range of the third confidence weight value are normalized to obtain the first confidence weight value, the second confidence weight value, and the third confidence weight value.

5. The method according to claim 1, wherein Before inputting the fused feature vector set into the abnormal behavior classification model, the method further includes: Acquire a training data set, where the training data set includes first sample data marked as normal behavior and second sample data marked as abnormal behavior; Inputting the first sample data and the second sample data into an initial abnormal behavior classification model for training to generate a trained abnormal behavior classification model; During the training process, a cross-validation method is used to evaluate the classification performance of the initial abnormal behavior classification model, and the hyperparameter configuration of the initial abnormal behavior classification model is adjusted according to the evaluation results.

6. The method according to claim 5, characterized in that The method further comprises: Obtaining original annotation information of each sample in the first sample data and the second sample data, and verifying the accuracy of the original annotation information; For each sample, extracting multimodal data corresponding to the sample based on the original annotation information, the multimodal data including sample visible light image data, sample infrared image data, and sample audio signal data; performing normalization processing on the sample visible light image data, the sample infrared image data, and the sample audio signal data respectively to generate standardized first sample data and second sample data; Divide the standardized first sample data and the second sample data into a training data set, a validation data set, and a test data set, and ensure that the distribution ratio of each type of sample in each data set is balanced during the division process; A preliminary evaluation of the performance of the initial abnormal behavior classification model is performed based on the validation dataset and the test dataset, and the sample weight configuration of the training dataset is adjusted according to the evaluation results.

7. The method according to any one of claims 2 to 6, characterized in that The method further comprises: During the target monitoring period, the environmental interference factor value in each monitoring period is calculated based on the real-time collection of the multimodal data; Determining the credibility level of the visible light image data, infrared image data, and audio signal data within each monitoring period based on the environmental interference factor value; The confidence weight values ​​of the first feature vector set, the second feature vector set, and the third feature vector set are dynamically adjusted according to the credibility level.

8. The method according to claim 7, characterized in that Calculating the environmental interference factor value in each monitoring period according to the real-time collection of the multimodal data during the target monitoring period includes the following steps: Obtain the light intensity, temperature, and noise intensity values ​​of the environment where the security camera is located during the target monitoring period; Calculate the comprehensive environmental parameters within each monitoring period based on the light intensity value, temperature value and noise intensity value; Based on the comprehensive environmental parameters and in combination with a preset environmental interference threshold range, determining the environmental interference factor value within each monitoring period; Determining the credibility level of the visible light image data, infrared image data, and audio signal data in each monitoring period based on the environmental interference factor value includes the following steps: For each monitoring period, a credibility level corresponding to the visible light image data, the infrared image data and the audio signal data is generated according to the environmental interference factor value and a preset weight distribution rule.

9. The method according to claim 8, characterized in that Calculating the comprehensive environmental parameters in each monitoring period according to the light intensity value, temperature value and noise intensity value comprises the following steps: Obtain the light intensity value sequence, temperature value sequence and noise intensity value sequence for each monitoring period within the target monitoring cycle; Performing smoothing filtering on the light intensity value sequence, temperature value sequence and noise intensity value sequence respectively to eliminate the influence of abnormal values; According to the smoothed light intensity value sequence, temperature value sequence and noise intensity value sequence, the mean light intensity, mean temperature and mean noise intensity in each monitoring period are calculated; Based on the light intensity mean, temperature mean and noise intensity mean, combined with preset weight coefficients, the comprehensive environmental parameters in each monitoring period are calculated.

10. A security camera abnormal behavior recognition system based on multimodal fusion, characterized in that: The system comprises: An acquisition module is used to acquire multimodal data collected by the security camera during a target monitoring period, wherein the multimodal data includes visible light image data, infrared image data, and audio signal data; a feature extraction module, configured to perform a first feature extraction process on the visible light image data to generate a first feature vector set, perform a second feature extraction process on the infrared image data to generate a second feature vector set, and perform a third feature extraction process on the audio signal data to generate a third feature vector set; a fusion module, configured to input the first feature vector set, the second feature vector set, and the third feature vector set into a multimodal feature fusion module, and generate a fused feature vector set by using a weighted fusion algorithm, wherein the weighted fusion algorithm dynamically adjusts the fusion weight according to the confidence weight of each modal feature vector; An abnormality classification module, configured to input the fused feature vector set into an abnormal behavior classification model to generate an abnormal behavior classification result, wherein the abnormal behavior classification result includes a normal behavior mark and an abnormal behavior mark; The alarm module is used to generate an abnormal behavior alarm signal according to the abnormal behavior classification result, and send the abnormal behavior alarm signal to the security monitoring terminal to trigger an alarm operation.

Citation Information

Cited By

  • Equipment operation and maintenance method and system based on multi-modal large model

    CN121052807A

  • Multi-mode home security anomaly detection method and device and medium

    CN121412848A

  • A multi-modal home security anomaly detection method, device and medium

    CN121412848B

  • Petrochemical engineering anomaly detection method and system based on multi-modal attention fusion

    CN121884065A

  • Petroleum chemical abnormality detection method and system based on multi-modal attention fusion

    CN121884065B