Video monitoring early warning method and system based on multi-mode behavior mode

By adopting the early warning method of multimodal behavior mode in the video surveillance system, multimodal data is extracted and integrated, and behavioral mode library and scene behavior baseline are established, which solves the problems of high false alarm rate and poor adaptability in the existing technology, and realizes automatic learning and dynamic adjustment of intelligent video surveillance early warning.

CN120219903AActive Publication Date: 2025-06-27SHENZHEN JOOAN TECH CO LTD

Patent Information

Application Number
CN202510696985.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing video surveillance system has high false alarm rate and poor adaptability. It is unable to dynamically learn behavior patterns in scenarios, and it is difficult to cope with changes in population density or new behavior types.

Method used

The video surveillance early warning method based on multimodal behavior mode is adopted, and multimodal data (RGB frames, optical flow fields, human skeleton key points, scene semantic segmentation diagrams) is extracted, and the behavior pattern library and scene behavior baseline are pre-processed and fusion are carried out, and the behavior mode library and the dynamic behavior mode library are updated based on online learning, and the abnormal judgment threshold is adjusted according to the current scene complexity.

Benefits of technology

It realizes an intelligent video surveillance early warning solution that automatically learns scene behavior patterns, dynamically adjusts warning thresholds, and integrates multi-dimensional feature analysis, which reduces the false alarm rate and improves adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219903A_ABST
    Figure CN120219903A_ABST
Patent Text Reader

Abstract

The invention provides a video monitoring early warning method based on a multi-modal behavior mode, and the method comprises the following steps: extracting multi-modal data according to a video stream, carrying out the preprocessing, carrying out the alignment and fusion of the preprocessed multi-modal data, and enabling the multi-modal data to comprise RGB frames, an optical flow field, human skeleton key points, and a scene semantic segmentation map; performing short-time behavior pattern feature and behavior pattern library establishment and scene behavior baseline construction on the fused multi-modal data; updating a dynamic behavior pattern library based on online learning, reducing a historical data weight in combination with a time decay factor, and adjusting an anomaly judgment threshold according to the complexity of a current scene; performing multi-level early warning according to the abnormal judgment threshold value, and performing feedback optimization; according to the method, an intelligent video monitoring early warning scheme which can automatically learn a scene behavior mode, dynamically adjust an early warning threshold value and fuse multi-dimensional feature analysis is provided, and the problems of high false alarm rate and poor adaptability in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent video surveillance, and particularly to a video surveillance warning method and system based on multi-modal behavior patterns. Background Art

[0002] Current video surveillance systems are mostly based on static rules (such as area intrusion and object left detection) or single behavior recognition (such as running and falling), and have the following defects: high false alarm rate: light changes, occlusion, and complex backgrounds are likely to trigger false alarms; poor adaptability: unable to dynamically learn behavior patterns in the scene, and difficult to cope with changes in crowd density or new behavior types; lack of correlation analysis: analyzing single frames or short-time segments in isolation, ignoring the spatio-temporal continuity of behaviors. Summary of the Invention

[0003] The main purpose of the present invention is to overcome the above defects in the prior art, and provide a video surveillance warning method and system based on multi-modal behavior patterns. This method provides an intelligent video surveillance warning solution that can automatically learn scene behavior patterns, dynamically adjust warning thresholds, and integrate multi-dimensional feature analysis, and solves the problems of high false alarm rate and poor adaptability in the prior art.

[0004] The present invention adopts the following technical solutions: A video surveillance warning method based on multi-modal behavior patterns includes the following steps: Extract multi-modal data from the video stream, perform preprocessing, and align and fuse the preprocessed multi-modal data. The multi-modal data includes RGB frames, optical flow fields, human skeleton key points, and scene semantic segmentation maps; Extract short-time behavior pattern features from the fused multi-modal data, establish a behavior pattern library, and construct a scene behavior baseline; Update the dynamic behavior pattern library based on online learning, combine a time decay factor to reduce the weight of historical data, and adjust the anomaly determination threshold according to the current scene complexity; Perform multi-level warnings according to the anomaly determination threshold and provide feedback for optimization.

[0005] Specifically, the alignment and fusion of the preprocessed multi-modal data specifically include: Strictly align the multi-modal data through hardware timestamps; Complement the intermediate frames of the low-frame-rate modality by linear interpolation; Superimpose the optical flow field on the RGB frame channels to obtain the first fusion data, and weight the human skeleton key point trajectories and the semantic segmentation map results through an attention mechanism to obtain the second fusion data.

[0006] Specifically, the extraction of short-time behavior pattern features from the fused multi-modal data, the establishment of a behavior pattern library, and the construction of a scene behavior baseline are specifically as follows: Extract the spatio-temporal features of the first fusion data and the second fusion data using a 3D CNN architecture, that is, construct a behavior pattern library for short-term behavior pattern features; Stitch the short-term behavior pattern features of the behavior pattern library according to a time window, and input a total of M vectors with a total span into a bidirectional LSTM to output a scene behavior baseline vector.

[0007] Specifically, the behavior pattern library is dynamically updated based on online learning, and the weight of historical data is reduced by combining a time decay factor. Specifically: Obtain short-term behavior pattern features in real time, and perform a sliding window incremental update on the behavior pattern library; And set the weights of the historical data in the behavior pattern library according to the generation time. The weight setting is specifically: ; Among them, is the initial weight, which is 1, is the decay rate, and the range is 0.001 - 0.01; is the interval between the current time and the data generation time.

[0008] Specifically, the anomaly detection threshold is adjusted according to the current scene complexity. Specifically: Extract the parameters in the current scene, including crowd density, light intensity, and movement activity, and perform normalization processing to obtain the normalized crowd density , light intensity ; Dynamic anomaly detection threshold : ; Among them, is the basic threshold, is the adjustment coefficient, and the higher the crowd density, the stronger the light intensity, and the larger the dynamic anomaly detection threshold.

[0009] Specifically, multi-level early warning is performed according to the anomaly detection threshold, and feedback optimization is performed. Specifically: Obtain the similarity score S(t) between the current behavior feature and the baseline; If S(t) < T(t): It is determined as abnormal and an early warning is entered; If S(t) ≥ T(t) but the administrator marks it as abnormal: perform feedback optimization and trigger incremental learning of the model.

[0010] On the other hand, the present invention also provides a multi-modal behavior pattern video surveillance early warning system, including: Data processing unit: Extract multimodal data from the video stream, perform preprocessing, and align and fuse the preprocessed multimodal data. The multimodal data includes RGB frames, optical flow fields, human skeleton key points, and scene semantic segmentation maps. Behavior baseline construction unit: Extract short-term behavior pattern features from the fused multimodal data, and establish a behavior pattern library and a scene behavior baseline. Update and adjustment unit: Update the dynamic behavior pattern library based on online learning, combine a time decay factor to reduce the weight of historical data, and adjust the anomaly detection threshold according to the current scene complexity. Early warning unit: Perform multi-level early warnings according to the anomaly detection threshold and provide feedback for optimization.

[0011] On the other hand, the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of a multimodal behavior pattern-based video surveillance early warning method are implemented.

[0012] On another aspect, the present invention provides a computer-readable storage medium, characterized in that a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of a multimodal behavior pattern-based video surveillance early warning method are implemented.

[0013] As can be seen from the above description of the present invention, compared with the prior art, the present invention has the following beneficial effects: The present invention proposes a multimodal behavior pattern-based video surveillance early warning method, including the following steps: Extract multimodal data from the video stream, perform preprocessing, and align and fuse the preprocessed multimodal data. The multimodal data includes RGB frames, optical flow fields, human skeleton key points, and scene semantic segmentation maps; Extract short-term behavior pattern features from the fused multimodal data, and establish a behavior pattern library and a scene behavior baseline; Update the dynamic behavior pattern library based on online learning, combine a time decay factor to reduce the weight of historical data, and adjust the anomaly detection threshold according to the current scene complexity; Perform multi-level early warnings according to the anomaly detection threshold and provide feedback for optimization. The method of the present invention provides an intelligent video surveillance early warning solution that can automatically learn scene behavior patterns, dynamically adjust the early warning threshold, and fuse multi-dimensional feature analysis, solving the problems of high false alarm rate and poor adaptability in the prior art. Description of the Drawings

[0014] Figure 1 It is a flowchart of a multimodal behavior pattern-based video surveillance early warning method provided by an embodiment of the present invention; Figure 2 is a structural diagram of a multimodal behavior pattern-based video surveillance early warning system provided by an embodiment of the present invention; FIG. 3 is a schematic diagram of an embodiment of an electronic device provided by an embodiment of the present invention; FIG. 4 is a schematic diagram of an embodiment of a computer-readable storage medium provided by an embodiment of the present invention. Detailed implementation manners

[0015] The present invention will be further described below through specific implementation manners.

[0016] The present invention proposes a multi-modal behavior pattern-based video surveillance and early warning method, which provides an intelligent video surveillance and early warning solution that can automatically learn scene behavior patterns, dynamically adjust early warning thresholds, and integrate multi-dimensional feature analysis, and solves the problems of high false alarm rate and poor adaptability in the prior art.

[0017] As Figure 1 is a multi-modal behavior pattern-based video surveillance and early warning method for the solution of the present invention; specifically includes the following steps: S101: Extract multi-modal data according to the video stream, perform preprocessing, and align and fuse the preprocessed multi-modal data. The multi-modal data includes RGB frames, optical flow fields, human body skeleton key points, and scene semantic segmentation maps; By fusing multi-dimensional visual information and dynamic environment parameters, problems such as high false detection rate and poor environmental adaptability caused by a single data source in traditional video surveillance are solved. The specific implementation is as follows: In the embodiment, a real-time video stream (supporting visible light / infrared dual-mode input, resolution ≥ 1080P, frame rate ≥ 25fps).

[0018] Extracted modality: RGB frames, the purpose is to obtain the original color information for object detection and scene semantic analysis. Extraction method: Intercept the video stream frame by frame and retain the timestamp alignment.

[0019] Optical flow field, the purpose is to quantify the pixel-level motion vector and capture short-term behavior dynamics (such as fast movement, limb swing). Algorithm selection: Farneback dense optical flow: Calculate the motion of all image pixels, suitable for complex scenes. Output format: Two-dimensional vector field (horizontal / vertical displacement matrix), with the same resolution as the original frame.

[0020] Human body skeleton key points, the purpose is to identify 18 key points of the human body (head, shoulders, elbows, wrists, etc.) and construct the pose spatio-temporal trajectory. Use multi-object tracking: Integrate the DeepSORT algorithm to solve the ID jump problem in the occlusion scene. Lightweight deployment: Use MobileNetV3 as the backbone network of OpenPose to achieve real-time inference at the edge (<30ms / frame).

[0021] Scene semantic segmentation map (Mask R-CNN), aiming to segment static elements (such as walls, railings) and dynamic objects (such as packages, vehicles) in the scene to assist in behavior context analysis.

[0022] Data preprocessing, preprocessing objectives: eliminate environmental interference and improve the robustness of subsequent feature extraction; Use the CLAHE algorithm for illumination normalization to solve the problem of target detection failure caused by sudden illumination changes (such as backlighting, shadows), convert the RGB frame to the LAB color space, and perform limited contrast adaptive histogram equalization only on the L channel (brightness).

[0023] Use the ViBe algorithm for background modeling to solve the problem of misjudging static objects (such as fixed facilities) as foreground targets; specifically including dynamic background update: each pixel maintains a background model containing 20 samples, randomly replacing historical samples (replacement probability 1 / 16), foreground detection: if the difference between the current pixel value and at least 2 samples in the background model is less than the threshold (set to 20), it is determined as the background.

[0024] Specifically, aligning and fusing the preprocessed multi-modal data specifically includes: Strictly align the multi-modal data through hardware timestamps; Use linear interpolation to complement the intermediate frames for the low-frame-rate modality; Overlay the optical flow field on the RGB frame channels to form a 5-channel input to obtain the first fusion data, (R / G / B / Flow_X / Flow_Y), for 3D CNN processing, and the second fusion data is obtained by weighting the human skeleton key point trajectory and the semantic segmentation map result through the attention mechanism.

[0025] Formula: ; Among them, σ is the Sigmoid function, and W1, W2 are learnable parameters.

[0026] S102: Extract short-term behavior pattern features from the fused multi-modal data, establish a behavior pattern library, and construct a scene behavior baseline; Specifically, extracting short-term behavior pattern features from the fused multi-modal data, establishing a behavior pattern library, and constructing a scene behavior baseline specifically include: Use the 3DCNN architecture to extract the spatio-temporal features of the first fusion data and the second fusion data, such as sudden acceleration, physical conflicts, which are short-term behavior pattern features, and construct a behavior pattern library; The first fusion data and the second fusion data, with a time series span of 5 seconds, 3D CNN (C3D network) architecture; this 3D CNN model is first trained, and the training data includes: Short-term behavior classification: Label definition: 20 predefined basic behaviors (such as "walking", "running", "raising hand", "throwing").

[0027] Training strategy: Pretraining: Transfer learning based on the Kinetics-400 dataset (400,000 video clips).

[0028] Fine-tuning: Use the target scene data (such as bank surveillance videos) to optimize the classification boundary.

[0029] Concatenate the short-term behavior pattern features of the behavior pattern library according to the time window. The total span is M vector inputs to the bidirectional LSTM, and the output is the baseline vector of the scene behavior, establishing a baseline to distinguish "statistical anomalies" from "real threats".

[0030] In the embodiment, the short-term behavior pattern features (4096 dimensions) are concatenated according to the time window (one vector every 5 seconds), and the total span is 30 minutes (a total of 360 vectors). Bidirectional LSTM: Stacked in 2 layers, with 256 hidden units in each layer to capture the temporal dependencies before and after.

[0031] S103: Update the dynamic behavior pattern library based on online learning, combine the time decay factor to reduce the weight of historical data, and adjust the anomaly determination threshold according to the current scene complexity; Specifically, update the behavior pattern library dynamically based on online learning, and combine the time decay factor to reduce the weight of historical data. Specifically: Obtain the short-term behavior pattern features in real time, and perform sliding window incremental update on the behavior pattern library; And set the weights for the historical data in the behavior pattern library according to the generation time. The weight setting is specifically: ; Among them, is the initial weight, which is 1, is the decay rate, and the range is 0.001 - 0.01; is the interval between the current time and the data generation time.

[0032] Specifically, and adjust the anomaly determination threshold according to the current scene complexity. Specifically: Extract the parameters in the current scene, including crowd density, light intensity, and movement activity, and perform normalization processing to obtain the normalized crowd density and light intensity ; Crowd density D(t): Count the number of dynamic targets / area of the region through the segmentation results of Mask R-CNN.

[0033] Light intensity L(t): Extract the mean value of the luminance channel (L value in the LAB color space) from the RGB frame.

[0034] Dynamic anomaly determination threshold : ; Among them, is the basic threshold, is the adjustment coefficient, and the higher the population density, the stronger the light intensity, and the larger the dynamic anomaly determination threshold.

[0035] S104: Perform multi-level early warning according to the anomaly determination threshold and feedback for optimization.

[0036] Specifically, performing multi-level early warning according to the anomaly determination threshold and feedback for optimization is specifically as follows: Obtain the similarity score S(t) between the current behavior feature and the baseline; the similarity is calculated as the cosine similarity between the current behavior feature and the baseline; If S(t) < T(t): Determine it as abnormal and enter the early warning; If S(t) ≥ T(t) but the administrator marks it as abnormal: Perform feedback optimization and trigger incremental learning of the model.

[0037] As shown in Figure 2, an embodiment of the present invention also provides a multi-modal behavior pattern video surveillance early warning system, including: Data processing unit 201: Extract multi-modal data according to the video stream, perform preprocessing, and align and fuse the preprocessed multi-modal data. The multi-modal data includes RGB frames, optical flow fields, human skeleton key points, and scene semantic segmentation maps; By fusing multi-dimensional visual information and dynamic environment parameters, problems such as high false detection rate and poor environmental adaptability caused by a single data source in traditional video surveillance are solved. The specific implementation is as follows: In the embodiment, a real-time video stream (supporting visible light / infrared dual-mode input, resolution ≥ 1080P, frame rate ≥ 25fps).

[0038] Extraction modality: RGB frame, the purpose is to obtain the original color information for object detection and scene semantic analysis. Extraction method: Intercept the video stream frame by frame and retain timestamp alignment.

[0039] Optical flow field, the purpose is to quantify pixel-level motion vectors and capture short-term behavior dynamics (such as fast movement, limb swing). Algorithm selection: Farneback dense optical flow: Calculate the motion of all image pixels, suitable for complex scenes. Output format: Two-dimensional vector field (horizontal / vertical displacement matrix), with the same resolution as the original frame.

[0040] Human skeletal key points, aiming to identify 18 key points of the human body (head, shoulders, elbows, wrists, etc.) and construct the spatio-temporal trajectory of the pose. Using multi-object tracking: integrating the DeepSORT algorithm to solve the problem of ID jumps in occlusion scenarios. Lightweight deployment: adopting MobileNetV3 as the backbone network of OpenPose to achieve real-time inference at the edge (<30ms / frame).

[0041] Scene semantic segmentation map (Mask R-CNN), aiming to segment static elements (such as walls, railings) and dynamic objects (such as packages, vehicles) in the scene to assist in behavior context analysis.

[0042] Data preprocessing, preprocessing objectives: eliminating environmental interference and enhancing the robustness of subsequent feature extraction; Using the illumination normalization CLAHE algorithm to solve the problem of target detection failure caused by sudden illumination changes (such as backlight, shadow), converting the RGB frame to the LAB color space, and only performing limited contrast adaptive histogram equalization on the L channel (brightness).

[0043] Using the background modeling ViBe algorithm to solve the problem of misjudging static objects (such as fixed facilities) as foreground targets; specifically including dynamic background update: maintaining a background model containing 20 samples for each pixel and randomly replacing historical samples (replacement probability 1 / 16), foreground detection: if the difference between the current pixel value and at least 2 samples in the background model is less than the threshold (set to 20), it is determined as the background.

[0044] Specifically, aligning and fusing the preprocessed multi-modal data specifically includes: Strictly aligning the multi-modal data through hardware timestamps; Completing the intermediate frames of the low-frame-rate modality by linear interpolation; Overlaying the optical flow field on the RGB frame channels to form a 5-channel input to obtain the first fusion data, (R / G / B / Flow_X / Flow_Y), for 3D CNN processing, and obtaining the second fusion data by weighting the human skeletal key point trajectory and the semantic segmentation map result through the attention mechanism.

[0045] Formula: ; where σ is the Sigmoid function, and W1, W2 are learnable parameters.

[0046] Behavior baseline construction unit 202: performing short-term behavior pattern features on the fused multi-modal data to establish a behavior pattern library and construct a scene behavior baseline; Specifically, performing short-term behavior pattern features on the fused multi-modal data to establish a behavior pattern library and construct a scene behavior baseline; specifically: The spatio-temporal features of the first fusion data and the second fusion data, such as sudden acceleration and physical conflicts, are extracted using a 3D CNN architecture, which are short-term behavior pattern features, and a behavior pattern library is constructed. The first fusion data and the second fusion data, with a time span of 5 seconds, are processed by a 3D CNN (C3D network) architecture; the 3D CNN model is first trained, and the training data includes: Short-term behavior classification: Label definition: 20 predefined basic behaviors (such as "walking", "running", "raising hands", "throwing") are defined.

[0047] Training strategy: Pre-training: Transfer learning based on the Kinetics-400 dataset (400,000 video clips).

[0048] Fine-tuning: Use the target scene data (such as bank surveillance videos) to optimize the classification boundary.

[0049] The short-term behavior pattern features in the behavior pattern library are concatenated according to a time window, and a total of M vectors are input into a bidirectional LSTM, and the output is the scene behavior baseline vector, and a baseline is established to distinguish "statistical anomalies" from "real threats".

[0050] In the embodiment, the short-term behavior pattern features (4096 dimensions) are concatenated according to a time window (one vector every 5 seconds), and the total span is 30 minutes (a total of 360 vectors). Bidirectional LSTM: Stacked in 2 layers, with 256 hidden units in each layer, capturing the temporal dependencies before and after.

[0051] Update and adjustment unit 203: Dynamically update the behavior pattern library based on online learning, combine a time decay factor to reduce the weight of historical data, and adjust the anomaly determination threshold according to the current scene complexity; Specifically, dynamically update the behavior pattern library based on online learning, and combine a time decay factor to reduce the weight of historical data, specifically: Obtain the short-term behavior pattern features in real time, and perform a sliding window-based incremental update on the behavior pattern library; And set weights for the historical data in the behavior pattern library according to the generation time, and the weight setting is specifically: ; Among them, is the initial weight, which is 1, is the decay rate, and the range is 0.001 - 0.01; is the interval between the current time and the data generation time.

[0052] Specifically, and adjust the anomaly determination threshold according to the current scene complexity, specifically: Extract the parameters in the current scene, including crowd density, light intensity, and movement activity, and perform normalization processing to obtain the normalized crowd density Light intensity ; Crowd density D(t): The number / area of dynamic objects is counted through the Mask R-CNN segmentation result.

[0053] Light intensity L(t): The mean value of the luminance channel (L value in the LAB color space) is extracted from the RGB frame.

[0054] Dynamic anomaly determination threshold : ; Among them, is the basic threshold, is the adjustment coefficient, and the higher the crowd density, the stronger the light intensity, and the larger the dynamic anomaly determination threshold.

[0055] Early warning unit 204: Perform multi-level early warning according to the anomaly determination threshold and feedback for optimization.

[0056] Specifically, perform multi-level early warning according to the anomaly determination threshold and feedback for optimization, specifically: Obtain the similarity score S(t) between the current behavior feature and the baseline; the similarity is calculated as the cosine similarity between the current behavior feature and the baseline; If S(t) < T(t): It is determined as an anomaly and enters the early warning; If S(t) ≥ T(t) but the administrator marks it as an anomaly: Perform feedback optimization and trigger model incremental learning.

[0057] As shown in Figure 3, an embodiment of the present invention provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, it implements a multi-modal behavior pattern video surveillance early warning method provided by an embodiment of the present invention.

[0058] In the specific implementation process, when the processor 320 executes the computer program 311, it can implement Figure 1 any implementation manner in the corresponding embodiment.

[0059] Since the electronic device introduced in this embodiment is the device adopted for implementing a data processing device in an embodiment of the present invention, based on the method introduced in an embodiment of the present invention, those skilled in the art can understand the specific implementation manner and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in an embodiment of the present invention will not be described in detail here. As long as the device adopted by those skilled in the art to implement the method in an embodiment of the present invention belongs to the scope of protection of the present invention.

[0060] Please refer to FIG. 4. FIG. 4 is a schematic diagram of an embodiment of a computer-readable storage medium provided by an embodiment of the present invention.

[0061] As shown in FIG. 4, this embodiment provides a computer-readable storage medium 400, on which a computer program 411 is stored. When the computer program 411 is executed by a processor, it implements a multi-modal behavior pattern video surveillance and early warning method provided by an embodiment of the present invention; In the specific implementation process, when the computer program 411 is executed by the processor, it can implement Figure 1 any one of the corresponding embodiments.

[0062] It should be noted that in the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0063] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0064] The present invention proposes a multi-modal behavior pattern video surveillance and early warning method, including the following steps: extracting multi-modal data from the video stream, performing preprocessing, and aligning and fusing the preprocessed multi-modal data. The multi-modal data includes RGB frames, optical flow fields, human skeleton key points, and scene semantic segmentation maps; performing short-term behavior pattern features on the fused multi-modal data, establishing a behavior pattern library and constructing a scene behavior baseline; updating the dynamic behavior pattern library based on online learning, combining a time decay factor to reduce the weight of historical data, and adjusting the anomaly determination threshold according to the current scene complexity; performing multi-level early warning according to the anomaly determination threshold and providing feedback for optimization. The method of the present invention provides an intelligent video surveillance and early warning solution that can automatically learn scene behavior patterns, dynamically adjust the early warning threshold, and fuse multi-dimensional feature analysis, solving the problems of high false alarm rate and poor adaptability in the prior art.

[0065] The above is only the specific implementation manner of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification made to the present invention using this concept shall fall within the scope of infringement of the protection of the present invention.

Claims

1. A multi-modal behavior pattern-based video surveillance and early warning method, characterized in that, It includes the following steps: Extract multimodal data from the video stream, perform preprocessing, and align and fuse the preprocessed multimodal data. The multimodal data includes RGB frames, optical flow fields, human skeleton key points, and scene semantic segmentation maps; Extract short-term behavior pattern features from the fused multimodal data, and establish a behavior pattern library and construct a scene behavior baseline; Update the dynamic behavior pattern library based on online learning, combine a time decay factor to reduce the weight of historical data, and adjust the anomaly determination threshold according to the current scene complexity; Perform multi-level early warnings according to the anomaly determination threshold and provide feedback for optimization.

2. The multimodal behavior pattern-based video surveillance and early warning method according to claim 1, wherein The alignment and fusion of the preprocessed multimodal data specifically include: Strictly align the multimodal data through hardware timestamps; Complement the intermediate frames of the low-frame-rate modality by linear interpolation; Superimpose the optical flow field on the RGB frame channel to obtain the first fused data, and obtain the second fused data by weighting the human skeleton key point trajectory and the semantic segmentation map result through an attention mechanism.

3. A multimodal behavior pattern-based video surveillance and early warning method according to claim 2, characterized in that The extraction of short-term behavior pattern features from the fused multimodal data, and the establishment of a behavior pattern library and the construction of a scene behavior baseline are specifically as follows: Use a 3D CNN architecture to extract the spatio-temporal features of the first fused data and the second fused data, which are the short-term behavior pattern features, and construct a behavior pattern library; Stitch the short-term behavior pattern features of the behavior pattern library according to a time window. The total span is M vectors, which are input into a bidirectional LSTM to output a scene behavior baseline vector.

4. A multimodal behavior pattern-based video surveillance and early warning method according to claim 2, characterized in that Dynamically update the behavior pattern library based on online learning, and combine a time decay factor to reduce the weight of historical data. Specifically: Obtain the short-term behavior pattern features in real time, and perform a sliding window-based incremental update on the behavior pattern library; Set the weights of the historical data in the behavior pattern library according to the generation time. The weight setting is specifically as follows: ; Among them, is the initial weight, which is 1, is the attenuation rate, and the range is 0.001 - 0.01; is the interval between the current time and the data generation time.

5. A multimodal behavior pattern-based video surveillance and early warning method according to claim 2, characterized in that, Adjust the anomaly determination threshold according to the current scene complexity. Specifically: Extract the parameters in the current scene, including crowd density, light intensity, and movement activity, and perform normalization to obtain the normalized crowd density , light intensity ; Dynamic exception determination threshold : ; Among them, is the basic threshold, is the adjustment coefficient, and the higher the population density, the stronger the light intensity, and the greater the dynamic anomaly determination threshold.

6. The multimodal behavior pattern-based video surveillance and early warning method according to claim 5, wherein, Perform multi-level early warnings according to the anomaly determination threshold and provide feedback for optimization. Specifically: Obtain the similarity score S(t) between the current behavior feature and the baseline; If S(t) < T(t): Determine it as an anomaly and enter the early warning; If S(t) ≥ T(t) but the administrator marks it as an anomaly: Provide feedback for optimization and trigger incremental learning of the model.

7. A multi-modal behavior pattern-based video surveillance and early warning system, characterized in that, It includes: A data processing unit: Extract multimodal data from the video stream, perform preprocessing, and align and fuse the preprocessed multimodal data. The multimodal data includes RGB frames, optical flow fields, human skeleton key points, and scene semantic segmentation maps; A behavior baseline construction unit: Extract short-term behavior pattern features from the fused multimodal data, and establish a behavior pattern library and construct a scene behavior baseline; An update and adjustment unit: Update the dynamic behavior pattern library based on online learning, combine a time decay factor to reduce the weight of historical data, and adjust the anomaly determination threshold according to the current scene complexity; An early warning unit: Perform multi-level early warnings according to the anomaly determination threshold and provide feedback for optimization.

8. An electronic device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the computer program, it implements the method steps described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Open domain video natural language description generation method based on multi-modal feature fusion

    CN108648746A

  • Multimedia system monitoring method and system based on data analysis

    CN116955092A

  • Intelligent access control management method and system based on multi-mode identification and Internet of Things technology

    CN118968665A

  • Video monitoring personnel behavior identification method and system

    CN119580352A

  • Elevator abnormal behavior detection system based on multi-mode neural network

    CN120024777A

Cited By

  • Space-time sequence data processing method and system for pet abnormal behavior recognition

    CN121071757A

  • Pet abnormal behavior recognition spatio-temporal sequence data processing method and system

    CN121071757B

  • Expressway scene video segmentation method and device and electronic equipment

    CN121600450A

  • A highway scene video segmentation method, device and electronic equipment

    CN121600450B

  • Operation maintenance personnel behavior analysis method and system of extra-high voltage transformer substation, electronic equipment and medium

    CN121765302A