Video analysis and behavior detection system and method based on self-learning mechanism

By employing a video analysis method based on a self-learning mechanism, the real-time performance and accuracy issues of traditional video surveillance systems in complex scenarios are addressed. This enables continuous analysis of target behavior between frames and abnormal early warning, thereby improving the real-time response capability of the video surveillance system.

CN120997758APending Publication Date: 2025-11-21青岛三三智能科技有限公司
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202511022109.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional video surveillance systems struggle to meet real-time and accuracy requirements in complex scenarios, particularly in terms of continuous analysis of inter-frame target behavior and anomaly warning.

Method used

A video analysis method based on a self-learning mechanism is adopted. By continuously extracting two frames of image data from the video stream, preprocessing and target detection are performed, the area to be detected is marked, the target behavior is analyzed, and an early warning is triggered.

Benefits of technology

It improves the real-time response capability of video surveillance, accurately identifies abnormal behavior and triggers early warnings in a timely manner, thus enhancing the real-time performance and accuracy of the video surveillance system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997758A_ABST
    Figure CN120997758A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video monitoring, in particular to a video analysis and behavior detection system and method based on a self-learning mechanism. The method comprises the following steps: continuously extracting two frames of image data from a video stream, preprocessing the image data, extracting a to-be-detected area of a current frame, and labeling a target in the to-be-detected area; according to the labeling result of the current frame, determining a to-be-detected area of the next frame, comparing the position and morphological change of the target in the two frames of images, and analyzing the behavior of the target; according to an analysis result, identifying an abnormal behavior of the target object, and triggering early warning; the system comprises a current frame target labeling module, a two-frame target change comparison module and an abnormal behavior recognition module. By extracting the to-be-detected area of the current frame, marking the target, linking the to-be-detected area of the next frame, carrying out target behavior analysis, identifying abnormal behaviors and triggering early warning, the real-time response capability of video monitoring is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video surveillance technology, and in particular to a video analysis and behavior detection system and method based on a self-learning mechanism. Background Technology

[0002] With the widespread application of video surveillance technology, traditional video surveillance systems mainly rely on manual monitoring, which suffers from problems such as low efficiency, slow response, and susceptibility to missed or false alarms. Especially in complex scenarios such as ports and transportation hubs, traditional methods struggle to meet the requirements of real-time performance and accuracy, and still fall short in terms of continuous analysis of target behavior across frames and anomaly warning.

[0003] Therefore, it is essential to propose a method to extract the detection region of the current frame and label the target, and then link it with the detection region of the next frame to perform target behavior analysis, thereby improving the real-time response capability of video surveillance. Summary of the Invention

[0004] The purpose of this invention is to provide a video analysis and behavior detection system and method based on a self-learning mechanism, aiming to solve the technical problems in the prior art that are difficult to meet the requirements of real-time performance and accuracy, and that there are still shortcomings in the continuous analysis of inter-frame target behavior and anomaly warning.

[0005] To achieve the above objectives, this invention employs a video analysis and behavior detection method based on a self-learning mechanism, comprising the following steps: Two frames of image data are extracted consecutively from the video stream. The image data is preprocessed, and the detection region of the current frame is extracted and the target in the detection region is marked. Based on the annotation results of the current frame, determine the detection area in the next frame, compare the position and shape changes of the target in the two frames, and analyze the target behavior; Based on the analysis results, abnormal behavior of the target object is identified and an alert is triggered.

[0006] Among the steps, in the process of extracting two consecutive frames of image data from the video stream, preprocessing the image data, extracting the detection region of the current frame, and labeling the target in the detection region: Read two frames of image data consecutively from a real-time video stream; A static detection region is preset, which is the region to be detected. Target detection is performed on the current frame for the region to be detected, and the target category and bounding box coordinates are output.

[0007] In the step of continuously reading two frames of image data from a real-time video stream: The image data is denoised and enhanced, and the image size is normalized.

[0008] In the preset static detection area, which is the area to be detected, target detection is performed on the current frame for the area to be detected, and the target category and bounding box coordinates are output: Draw bounding boxes on the image and label them with category tags.

[0009] Among the steps, drawing bounding boxes on the image and labeling them with category tags: Record the target feature vector.

[0010] Among them, in the steps of determining the detection region in the next frame based on the annotation results of the current frame, comparing the position and shape changes of the target in the two frames, and analyzing the target behavior: Predict the target location in the next frame based on the target bounding box in the current frame; Extract feature points of targets in two frames, match the same target, compare changes in the bounding boxes of targets in the two frames, and identify abnormal target shapes; Based on the abnormality of the target's form, determine the abnormality of the target's behavior.

[0011] In the step of predicting the target position in the next frame based on the target bounding box in the current frame: The detection range of the next frame image is narrowed, and a pre-defined electronic fence area is used to perform behavior analysis on targets within the fence.

[0012] In the step of identifying abnormal behavior of the target object based on the analysis results and triggering an early warning: When the target behavior matches the preset rules, it is immediately marked as an anomaly, and the target behavior sequence is analyzed to identify complex anomaly patterns; Trigger a multi-level early warning mechanism and push out early warning messages.

[0013] After triggering the multi-level early warning mechanism and pushing out early warning messages: Video clips can be retrieved based on alarm timestamps to trace back the events.

[0014] This invention also provides a video analysis and behavior detection system based on a self-learning mechanism, including a current frame target annotation module, a two-frame target change comparison module, and an abnormal behavior recognition module; wherein: The current frame target annotation module is used to continuously extract two frames of image data from the video stream, preprocess the image data, extract the detection area of ​​the current frame, and annotate the targets in the detection area. The two-frame target change comparison module is used to determine the detection area in the next frame based on the annotation results of the current frame, compare the position and shape changes of the target in the two frames, and analyze the target behavior. The abnormal behavior recognition module is used to identify abnormal behavior of the target object based on the analysis results and trigger an early warning.

[0015] This invention discloses a video analysis and behavior detection system and method based on a self-learning mechanism. The system employs a current frame target annotation module, a two-frame target change comparison module, and an abnormal behavior recognition module to perform the following steps: continuously extracting two frames of image data from the video stream; preprocessing the image data; extracting the detection region of the current frame and annotating the targets within that region; determining the detection region of the next frame based on the annotation results of the current frame; comparing the position and shape changes of the targets in the two frames to analyze target behavior; identifying abnormal behavior of the target object based on the analysis results and triggering an early warning; and improving the real-time response capability of video surveillance by extracting the detection region of the current frame, annotating the targets, and linking it with the detection region of the next frame. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of the steps of the video analysis and behavior detection method based on the self-learning mechanism of the present invention.

[0018] Figure 2 This is a flowchart of steps S100 of the present invention.

[0019] Figure 3 This is a flowchart of steps S200 of the present invention.

[0020] Figure 4 This is a flowchart of steps S300 of the present invention.

[0021] Figure 5 This is a schematic diagram of the video analysis and behavior detection system based on a self-learning mechanism according to the present invention.

[0022] Figure 6 This is a schematic diagram of the electronic device of the present invention.

[0023] 401 - Current frame target annotation module, 402 - Two-frame target change comparison module, 403 - Abnormal behavior recognition module. Detailed Implementation

[0024] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.

[0025] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0026] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0027] Please see Figures 1-4 This invention provides a video analysis and behavior detection method based on a self-learning mechanism, comprising the following steps: S100: Extract two consecutive frames of image data from the video stream, preprocess the image data, extract the detection area of ​​the current frame, and label the target in the detection area.

[0028] In this embodiment, two frames of image data are extracted consecutively from the video stream, the image data is preprocessed, and the detection region of the current frame is extracted and the target in the detection region is labeled. The specific process is as follows: S101: Read two frames of image data continuously from the real-time video stream, denoise and enhance the image data, and normalize the image size; S102: Preset static detection area, which is the area to be detected. Target detection is performed on the current frame for the area to be detected, and the target category and bounding box coordinates are output. S103: Draw bounding boxes on the image and label them with category tags, while recording the target feature vectors.

[0029] In the above process, frame extraction involves continuously reading two frames of image data (the current frame F) from a real-time video stream (such as an RTSP protocol camera). t and the next frame F t+1A circular buffer mechanism is used to ensure a stable inter-frame time interval (e.g., 33ms@30FPS).

[0030] Denoising and Enhancement: Non-local means denoising algorithms are used to eliminate image noise (such as noise caused by camera shake). Contrast in low-light areas is improved through CLAHE (Contrast-Limited Adaptive Histogram Equalization) to adapt to nighttime operation scenarios.

[0031] Size normalization: Images are uniformly scaled to 640×640 pixels (YOLOv5 input size), and bilinear interpolation is used to maintain image quality.

[0032] Static detection area preset and target detection, area division: Preset static detection areas (such as gantry crane track, grab bucket operation area) according to scene requirements, and draw rectangular areas on the image.

[0033] Object detection: Load a pre-trained YOLOv5s model. Input the current frame F t Output the target category (e.g., "Person", "Vehicle", "Cargo") and bounding box coordinates (in [x... min y min x max y max ]).

[0034] Target labeling and feature recording, visual labeling: Draw bounding boxes on the image (color distinguishes categories, such as red for people and blue for cars) and label the categories (such as "Person:ID001").

[0035] Feature extraction: Extract feature vectors from detected targets (e.g., ArcFace for faces, license plate OCR results for vehicles).

[0036] The feature vectors are stored in an in-memory database (Redis) for subsequent inter-frame target matching.

[0037] Two frames of image data are continuously read from a real-time video stream. The image data is denoised and enhanced, and the image size is normalized. A static detection region is preset, which is the region to be detected. Target detection is performed on the current frame for the region to be detected, and the target category and bounding box coordinates are output. Bounding boxes are drawn on the image and labeled with category labels, while the target feature vector is recorded. By acquiring high-quality image data from the video stream, the region to be detected in the current frame is accurately located and the target is labeled, providing a foundation for subsequent inter-frame analysis.

[0038] S200: Based on the annotation results of the current frame, determine the detection area of ​​the next frame, compare the position and shape changes of the target in the two frames, and analyze the target behavior.

[0039] In this embodiment, based on the annotation results of the current frame, the region to be detected in the next frame is determined. The position and morphological changes of the target in the two frames are compared to analyze the target behavior. The specific process is as follows: S201: Predict the target position in the next frame based on the target bounding box in the current frame, narrow the detection range of the next frame image, and preset the electronic fence area to perform behavior analysis on the target within the fence. S202: Extract feature points of targets in two frames, match the same target, compare changes in the bounding boxes of targets in two frames, and identify abnormal target shapes; S203: Determine abnormal target behavior based on abnormal target morphology.

[0040] In the above process, motion prediction involves applying a Kalman filter to the target bounding box in the current frame to predict the coordinates (x, y) of the target center point in the next frame. t+1 y t+1 Expand the detection area based on the predicted location (e.g., expand the bounding box by 20%) to reduce computational load.

[0041] Load the preset geofence coordinates (such as the vertices of a polygon in the track area) and determine whether the target has entered the fence. Perform subsequent behavior analysis only on targets inside the fence.

[0042] Inter-frame target matching and morphological anomaly detection, feature point matching: extract feature points of targets in two frames, and calculate the correspondence of feature points through the FLANN matcher.

[0043] Bounding box change analysis: Calculate the IoU (Intersection over Union) of the bounding boxes of two frames. If the IoU < 0.3, it is determined that the target is lost or a new target has appeared. Analyze the bounding box size change rate (e.g., a height change > 50% may indicate that a person has fallen).

[0044] Abnormal behavior determination, rule engine: preset behavior rules (such as "personnel stay in the track for more than 5 seconds"), and track the target behavior sequence through state machine.

[0045] If the rule is triggered, it is marked as a primary exception (e.g., "Potential Trespassing").

[0046] Temporal model analysis: Input the target location sequence into the LSTM network to identify complex abnormal patterns (such as theft behavior of "loitering → approaching the device → operating").

[0047] Model adaptive optimization and parameter adjustment: The confidence threshold of YOLOv5 is dynamically adjusted according to the false detection rate (e.g., the threshold is increased from 0.5 to 0.7 when there is a false detection).

[0048] Optimize the IoU threshold of NMS (e.g., adjust it from 0.45 to 0.5) to reduce overlapping bounding boxes.

[0049] Anomaly sample library update: Newly discovered video clips of abnormal behavior (such as violations) are labeled and stored in HDFS.

[0050] Predict the target position in the next frame based on the target bounding box in the current frame, narrow the detection range of the next frame image, and preset the electronic fence area to perform behavior analysis on the target within the fence; extract feature points of the target in two frames, match the same target, and compare the changes in the target bounding boxes in the two frames to identify target morphological anomalies; judge the target behavior anomalies based on the target morphological anomalies; identify target morphological / positional anomalies through inter-frame target correlation and behavior analysis.

[0051] S300: Identify abnormal behavior of the target object based on the analysis results and trigger an alert.

[0052] In this embodiment, abnormal behavior of the target object is identified based on the analysis results, and an early warning is triggered. The specific process is as follows: S301: When the target behavior conforms to the preset rules, it is immediately marked as abnormal, and the target behavior sequence is analyzed to identify complex abnormal patterns; S302: Triggers a multi-level early warning mechanism and pushes an early warning message; S303: Retrieve video clips based on alarm timestamps to perform event backtracking.

[0053] In the above process, abnormal behavior identification and rule-based model fusion judgment are performed: If either the rule engine or the LSTM model determines an anomaly, an alert will be triggered.

[0054] Record the anomaly type (such as "Trespassing" or "FallDetection") and the confidence score.

[0055] Multi-level warnings triggered, on-site alarms: A red alert window pops up on the monitoring interface, displaying screenshots of the abnormal event and handling suggestions (such as "immediately remove personnel").

[0056] Trigger the camera flash and buzzer (controlled via GPIO interface).

[0057] Remote notification: Send alert messages (including links to event videos) to the administrator's mobile app.

[0058] Equipment linkage: A stop command is sent to the PLC via the Modbus TCP protocol (e.g., when someone is detected under the grab bucket).

[0059] Event recap, video search: Retrieve video clips (H.265 encoded) from HDFS within 10 seconds before and after the alarm timestamp.

[0060] Use FFmpeg to generate thumbnail timelines for quick keyframe location.

[0061] Structured storage: Store the event metadata (target ID, behavior type, time, location) in a MySQL database.

[0062] Video clips are stored in hierarchical order by "year / month / day".

[0063] When the target behavior meets the preset rules, it is immediately marked as abnormal, and the target behavior sequence is analyzed to identify complex abnormal patterns; a multi-level early warning mechanism is triggered and early warning messages are pushed; video clips are retrieved based on the alarm timestamp to perform event backtracking; multi-level real-time early warning is realized, and a complete event audit chain is provided.

[0064] Corresponding to the aforementioned embodiments of video analysis and behavior detection methods based on self-learning mechanisms, this application also provides embodiments of video analysis and behavior detection systems based on self-learning mechanisms.

[0065] Figure 5 This is a block diagram of a video analysis and behavior detection system based on a self-learning mechanism, according to an exemplary embodiment. (Refer to...) Figure 5 The system may include: a current frame target annotation module 401, a two-frame target change comparison module 402, and an abnormal behavior recognition module 403; wherein: The current frame target annotation module 401 is used to continuously extract two frames of image data from the video stream, preprocess the image data, extract the detection area of ​​the current frame, and annotate the target in the detection area. The two-frame target change comparison module 402 is used to determine the detection area of ​​the next frame based on the annotation result of the current frame, compare the position and shape changes of the target in the two frames of images, and analyze the target behavior. The abnormal behavior recognition module 403 is used to identify abnormal behavior of the target object based on the analysis results and trigger an early warning.

[0066] In this embodiment, the current frame target annotation module 401 continuously extracts two frames of image data from the video stream, preprocesses the image data, extracts the detection area of ​​the current frame, and annotates the target in the detection area; the two-frame target change comparison module 402 determines the detection area of ​​the next frame based on the annotation result of the current frame, compares the position and shape changes of the target in the two frames, and analyzes the target behavior; the abnormal behavior recognition module 403 identifies the abnormal behavior of the target object based on the analysis result and triggers an early warning; by extracting the detection area of ​​the current frame, annotating the target, and linking the detection area of ​​the next frame, target behavior analysis is performed, abnormal behavior is identified, and an early warning is triggered, thereby improving the real-time response capability of video surveillance.

[0067] Regarding the system in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0068] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0069] Accordingly, this application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the video analysis and behavior detection method based on the self-learning mechanism described above. Figure 6 The diagram shown is a hardware structure diagram of any data processing device in which a video analysis and behavior detection system based on a self-learning mechanism, provided by an embodiment of the present invention, is located. (Except for...) Figure 6 In addition to the processor, memory, and network interface shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.

[0070] Accordingly, this application also provides a computer-readable storage medium storing computer instructions thereon, which, when executed by a processor, implement the video analysis and behavior detection method based on the self-learning mechanism described above. The computer-readable storage medium can be an internal storage unit of any data processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.

[0071] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0072] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A video analysis and behavior detection method based on a self-learning mechanism, characterized in that, Includes the following steps: Two frames of image data are extracted consecutively from the video stream. The image data is preprocessed, and the detection region of the current frame is extracted and the target in the detection region is labeled. Based on the annotation results of the current frame, determine the detection area in the next frame, compare the position and shape changes of the target in the two frames, and analyze the target behavior; Based on the analysis results, abnormal behavior of the target object is identified and an alert is triggered.

2. The video analysis and behavior detection method based on a self-learning mechanism as described in claim 1, characterized in that, In the steps of extracting two consecutive frames of image data from a video stream, preprocessing the image data, extracting the detection region of the current frame, and labeling the target in the detection region: Read two frames of image data consecutively from a real-time video stream; A static detection region is preset, which is the region to be detected. Target detection is performed on the current frame for the region to be detected, and the target category and bounding box coordinates are output.

3. The video analysis and behavior detection method based on a self-learning mechanism as described in claim 2, characterized in that, In the step of continuously reading two frames of image data from a real-time video stream: The image data is denoised and enhanced, and the image size is normalized.

4. The video analysis and behavior detection method based on a self-learning mechanism as described in claim 2, characterized in that, After the steps of performing target detection on the current frame within the preset static detection area (the area to be detected) and outputting the target category and bounding box coordinates: Draw bounding boxes on the image and label them with category tags.

5. The video analysis and behavior detection method based on a self-learning mechanism as described in claim 4, characterized in that, In the steps of drawing bounding boxes on an image and labeling them with category tags: Record the target feature vector.

6. The video analysis and behavior detection method based on a self-learning mechanism as described in claim 1, characterized in that, In the steps of determining the detection region for the next frame based on the annotation results of the current frame, comparing the position and shape changes of the target in the two frames, and analyzing the target behavior: Predict the target location in the next frame based on the target bounding box in the current frame; Extract feature points of targets in two frames, match the same target, compare changes in the bounding boxes of targets in the two frames, and identify abnormal target shapes; Based on the abnormality of the target's form, determine the abnormality of the target's behavior.

7. The video analysis and behavior detection method based on a self-learning mechanism as described in claim 6, characterized in that, In the step of predicting the target location in the next frame based on the target bounding box in the current frame: The detection range of the next frame image is narrowed, and a pre-defined electronic fence area is used to perform behavior analysis on targets within the fence.

8. The video analysis and behavior detection method based on a self-learning mechanism as described in claim 1, characterized in that, In the step of identifying abnormal behavior of the target object based on the analysis results and triggering an alert: When the target behavior matches the preset rules, it is immediately marked as an anomaly, and the target behavior sequence is analyzed to identify complex anomaly patterns; Trigger a multi-level early warning mechanism and push out early warning messages.

9. The video analysis and behavior detection method based on a self-learning mechanism as described in claim 8, characterized in that, After triggering the multi-level early warning mechanism and pushing out early warning messages: Video clips can be retrieved based on alarm timestamps to trace back the events.

10. A video analysis and behavior detection system based on a self-learning mechanism, applied to the video analysis and behavior detection method based on a self-learning mechanism as described in claim 1, characterized in that, This includes a current frame target annotation module, a two-frame target change comparison module, and an abnormal behavior recognition module; among which: The current frame target annotation module is used to continuously extract two frames of image data from the video stream, preprocess the image data, extract the detection area of ​​the current frame, and annotate the targets in the detection area. The two-frame target change comparison module is used to determine the detection area in the next frame based on the annotation results of the current frame, compare the position and shape changes of the target in the two frames, and analyze the target behavior. The abnormal behavior recognition module is used to identify abnormal behavior of the target object based on the analysis results and trigger an early warning.

Citation Information

Patent Citations

  • Comparison network-based target tracking method and device

    CN106920247A

  • Data analysis method of big data platform

    CN107480618A

  • Target object detection method, device and system, and neural network structure

    CN108073864A

  • Video analyzing-processing method and device

    CN108416301A

  • Method of detecting video static target

    CN110348394A