Abnormal behavior person identification method based on multi-source monitoring image integration

Through the multi-source surveillance image integration method, multi-camera collaborative acquisition and posture estimation, and semantic segmentation model, a cross-view behavior trajectory sequence is constructed, which solves the accuracy and completeness problems of abnormal behavior identification under a single perspective and achieves a deep understanding of behavioral intentions and trajectories.

CN120689814AActive Publication Date: 2025-09-23LIAOCHENG TIANYUAN ELECTRONIC ENG CO LTD

Patent Information

Application Number
CN202510810338.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-23
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

In the existing technology, the video image recognition model based on a single perspective has problems such as different monitoring point deployment angles, complex personnel behavior paths, severe occlusion, and difficulty in cross-perspective identity association in the monitoring scene. As a result, the abnormal behavior recognition results are susceptible to interference, and there is a lack of coordinated analysis of behavior trajectories, regional rules, and behavior deviation patterns. It is difficult to support the needs of public security in actual operations to trace the behavior trajectories and identify intentions of personnel.

Method used

A multi-source surveillance image integration method is adopted to collaboratively collect video data from different perspectives and time periods through multiple cameras. Combined with the posture estimation model and the human semantic segmentation model, high-correlation image frame sequences are extracted, and a cross-perspective personnel behavior trajectory sequence is constructed. Abnormal behavior deviation analysis is also performed, including behavior trajectory encoding, preset template matching and confidence aggregation value calculation.

Benefits of technology

It achieves precise association of identities across perspectives, ensures the integrity and accuracy of behavioral expression, improves the accuracy and explainability of abnormal behavior identification, breaks through the limitations of a single perspective, and supports deep semantic analysis of behavioral intentions and trajectories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689814A_ABST
    Figure CN120689814A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of abnormal behavior person recognition, and particularly relates to an abnormal behavior person recognition method based on multi-source monitoring image integration. The method comprises the steps of obtaining original video data of a plurality of cameras in a target area; judging whether a person completely appears or partially disappears through an attitude estimation and semantic segmentation model, carrying out inter-frame sampling by taking the judgment as a starting point and a stopping point, and retaining time-space information; preprocessing the image frames, and extracting static posture and dynamic behavior characteristics; fusing dynamic behaviors, appearance and position information to realize cross-camera personnel identity association and construct a behavior track sequence; and matching the trajectory with a preset template, and calculating a deviation score in combination with time consistency, position path similarity, an action matching degree and an abnormal fragment confidence aggregation value to recognize abnormal personnel. The method breaks through the limitation of a single visual angle, improves the low-illumination recognition precision, solves the problem of cross-visual-angle identity breakage, and improves the recognition accuracy and traceability of abnormal behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of abnormal behavior person identification, and in particular relates to an abnormal behavior person identification method based on multi-source monitoring image integration. Background Art

[0002] With the accelerated development of smart cities, public security, and public safety control systems, intelligent behavior recognition technology based on video surveillance has become a key technical tool in the practical application of public security video imagery. Existing technologies typically use single-view video image recognition models to identify behaviors and provide anomaly warnings for individuals appearing in surveillance footage. However, in real-world scenarios, issues such as varying monitoring point angles, complex individual behavior paths, significant occlusion, and difficulty in cross-viewpoint identity association make abnormal behavior recognition results susceptible to interference, leading to the risk of missed and misidentified behaviors. Furthermore, some behaviors exhibit phased and chained characteristics, making it difficult to reconstruct behavioral intent based solely on a single frame or perspective. This is especially true in areas covered by multiple cameras. Inconsistent recognition results for the same individual from different cameras further reduce the accuracy of the system's assessment of behavioral risk. Furthermore, traditional abnormal behavior recognition methods often rely on the output of action classification models and lack integrated semantic analysis of behavioral trajectories, regional rules, and behavioral deviation patterns. This makes it difficult to support the practical needs of public security for tracing individual behavioral trajectories, reconstructing behavioral chains, and identifying intent. Summary of the Invention

[0003] In view of the technical problems existing in the above background technology, the present invention proposes a method for identifying people with abnormal behavior by integrating multi-source monitoring images.

[0004] In order to achieve the above object, the technical solution adopted by the present invention includes the following steps:

[0005] S1. Video data acquisition: acquiring raw video data from multiple surveillance cameras deployed in the target surveillance area, wherein the video data includes image streams from different viewing angles and different time periods;

[0006] S2. Video image frame extraction: performing inter-frame sampling on the original video stream, extracting an image frame sequence, and retaining the timestamp, camera number, and location information corresponding to each frame;

[0007] S3, image preprocessing and feature extraction: preprocessing the image frame sequence and extracting the behavioral features of the human target in the image frame sequence;

[0008] S4. Personnel trajectory reconstruction and identity association: The same person identified by different cameras is matched in time and space to construct a cross-viewpoint sequence of person behavior trajectories;

[0009] S5. Behavioral rule deviation analysis: Match the person's behavior trajectory with the preset template and calculate the abnormal behavior deviation score. If the abnormal behavior deviation score exceeds the set deviation score threshold, the person is identified as an abnormal behavior person.

[0010] Preferably, the operation of performing inter-frame sampling on the original video stream in step S2 is to first determine that a person appears completely in the video stream as the start of inter-frame sampling, and to determine that a person disappears partially as the end of inter-frame sampling.

[0011] Preferably, the method for judging complete appearance and partial disappearance is:

[0012] Key point network acquisition: Using posture estimation model to collect key nodes of the human body in real time;

[0013] Preliminary complete appearance judgment: When all the key nodes collected are located in the valid area of ​​the image and the confidence of the key points is higher than the set confidence threshold, it is judged to be preliminarily complete appearance;

[0014] Preliminary partial disappearance judgment: define the importance weights of key nodes and calculate the weighted ratio of disappeared nodes. When the disappearance ratio exceeds the set threshold, it is judged as preliminary partial disappearance;

[0015] Enhanced judgment effect: Use the human body semantic segmentation model to segment the pixel area of ​​the human body and calculate the coverage rate of the human body mask and the image area. When the coverage rate is ≥90%, the auxiliary confirmation will judge the initial complete appearance as complete appearance; when the coverage rate is ≤40%, the auxiliary confirmation will judge the initial partial disappearance as partial disappearance.

[0016] Preferably, the image preprocessing and feature extraction operation in step S3 includes:

[0017] Performing image enhancement processing on the image frame, wherein the image enhancement processing includes image contrast enhancement, illumination normalization processing and edge denoising processing, so as to improve the recognition accuracy of human body areas in low-light environments;

[0018] Using the key node information of the human body in each frame of image, a static posture feature vector is constructed;

[0019] The dynamic behavior characteristics of the human body are further extracted by combining the change trend of key points between adjacent image frames to describe the action state of the human body in the frame image.

[0020] Preferably, the operation of reconstructing the personnel trajectory and associating it with the identity in step S4 includes:

[0021] Utilize the dynamic behavior characteristics of the human body, and introduce the appearance features and position information for fusion to form a joint feature representation for cross-camera identity matching;

[0022] In multiple camera image frame sequences, the identities of people appearing in different cameras are associated based on joint feature representation to determine whether they belong to the same target individual;

[0023] Based on the successful identity association, a sequence of people's behavior trajectories is constructed according to the time sequence, spatial position and behavioral actions of the people appearing under different cameras.

[0024] Preferably, the specific operations of matching the personnel behavior trajectory with the preset template and calculating the abnormal behavior deviation score in step S5 include:

[0025] Encode the behavior trajectory of the person into a behavior sequence sorted by time, wherein the behavior sequence consists of a location node, a behavior action, and a behavior time;

[0026] Select the normal behavior template sequence that is closest to the person's behavior type and path structure, and calculate the trajectory matching score S based on time consistency, location path similarity and action sequence matching degree. match ; The calculation method of time consistency is Where N is the trajectory length, T t Represents the timestamp of time t, T t norm Represents the timestamp of template time t; the calculation method of location path similarity is Where M represents the total number of frames of the normal template, L t Represents the position feature at time t, Represents the position feature of the template at time t; the calculation method of the action sequence matching degree is: Where K represents the number of key action points, I() is the indicator function, which returns 1 if the condition is met and returns 0 if the condition is satisfied. Represents the behavioral action characteristics at time t, represents the action feature of the template at time t; the trajectory matching score is obtained by weighted fusion of the three;

[0027] Identify multiple potential abnormal behavior segments in the behavior trajectory and assign an abnormal behavior confidence C to each abnormal behavior segment. j , combined with the duration to calculate the abnormal behavior confidence aggregation value S abn ; The confidence aggregate value S abn The calculation method is: Where J represents the number of identified potential abnormal behavior segments;

[0028] Finally, the trajectory matching score and the anomaly aggregation value are combined to construct the final deviation score S deviation , calculated as: S deviation =(1-S match )·Sabn .

[0029] Compared with the prior art, the advantages and positive effects of the present invention are:

[0030] 1. Multiple cameras are used to collaboratively collect video data from different perspectives and time periods, breaking through the limitations of single-perspective information. By combining a pose estimation model with a human semantic segmentation model, sampling starts when a person fully appears and ends when they partially disappear. Highly correlated image frame sequences are dynamically extracted, avoiding the redundant interference of traditional fixed-frame-rate sampling and ensuring the integrity of behavioral expression.

[0031] 2. A joint feature representation for cross-camera identity matching is constructed by integrating dynamic behavioral features, appearance features, and location information. Based on spatiotemporal constraints and a bidirectional matching consistency mechanism, this allows for accurate identification of the same person across different cameras, constructing a continuous cross-view behavioral trajectory sequence and resolving the cross-view identity fragmentation issue in existing technologies.

[0032] 3. Encode the person's behavior trajectory into a sequence of location, action, and time, perform a multi-dimensional match against a pre-set normal template, and incorporate confidence aggregation values ​​for abnormal behavior segments to comprehensively calculate the final deviation score. This breaks through the one-sidedness of traditional single-action classification and enables in-depth semantic analysis of behavioral intent and trajectory deviation patterns, improving the accuracy and interpretability of abnormal behavior identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0034] Figure 1 This is a flowchart of the overall structure of a method for identifying people with abnormal behavior by integrating multi-source surveillance images. DETAILED DESCRIPTION

[0035] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described below in conjunction with the accompanying drawings and embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.

[0036] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways than those described herein. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0037] Embodiment: In the field of smart cities and public security, abnormal behavior recognition based on video surveillance is a core technology for preventing security risks. Existing methods generally rely on video image models from a single perspective, which makes it difficult to cope with multiple challenges in complex scenarios. First, in order to solve the information limitations and video data redundancy problems of a single perspective and to achieve complete capture of multi-dimensional behavioral characteristics, a multi-camera collaborative acquisition and intelligent frame sampling solution is adopted. Video data acquisition: The original video data of multiple surveillance cameras deployed in the target monitoring area is acquired, and the video data contains image streams from different perspectives and different time periods. Specifically, multi-angle cameras are deployed in the target monitoring area to synchronously collect original video streams from different perspectives and time periods, covering multi-scene actions such as people moving, staying, and turning.

[0038] Then, video image frame extraction is performed: inter-frame sampling is performed on the original video stream to extract a sequence of image frames, retaining the timestamp, camera number, and location information corresponding to each frame. Inter-frame sampling of the original video stream begins by first determining the complete presence of a person in the video stream, and ends by determining their partial disappearance. The judgment of complete presence and partial disappearance is performed as follows: key point network acquisition: using a posture estimation model to collect key human nodes in real time; preliminary complete presence judgment: when all collected key nodes are within the valid image area and the confidence of the key points exceeds a set confidence threshold, a preliminary complete presence is determined; preliminary partial disappearance judgment: defining key node importance weights and calculating the weighted proportion of disappeared nodes. When the disappearance proportion exceeds a set threshold, a preliminary partial disappearance is determined; and judgment enhancement: using a human semantic segmentation model to segment the pixel area of ​​the human body, calculating the coverage ratio between the human mask and the image area, when the coverage ratio is ≥90%, auxiliary confirmation judges preliminary complete presence as complete presence; when the coverage ratio is ≤40%, auxiliary confirmation judges preliminary partial disappearance as partial disappearance. Specifically, a pose estimation model scans image frames in real time, locates key human nodes, and outputs the two-dimensional coordinates and confidence scores for each node. For example, when a person is standing straight ahead, the model fully captures the head, torso, and limb nodes. If the person is obscured from the side, the confidence scores of some nodes may fall below a threshold. The model then determines whether all key nodes are within the valid image area. A confidence threshold is set, and only when the confidence scores of all nodes exceed this threshold is the person considered to be preliminarily fully present. Preliminary partial disappearance is determined by first assigning importance weights to key nodes (e.g., 0.3 for the head, 0.2 for the torso, 0.15 for the upper limbs, and 0.15 for the lower limbs). Core nodes (e.g., the head and torso) are given higher weights to reflect their impact on action recognition. The disappearance ratio is then calculated by summing the weights of the nodes not detected in the current frame and dividing this by the total weight to obtain the weighted proportion of disappeared nodes. When this proportion exceeds the set threshold, the person is considered to be preliminarily partially disappeared. This is followed by an enhanced judgment: a person mask is generated using the human semantic segmentation model, and the coverage ratio of the mask pixels to the image area is calculated. Complete appearance confirmation: If the coverage rate is ≥90%, it indicates that the human body is completely in the frame, and the auxiliary confirmation of the preliminary complete appearance is complete appearance, eliminating the situation where the posture estimation is misjudged due to occlusion. Partial disappearance confirmation: If the coverage rate is ≤40%, it indicates that most of the human body has left the picture, and the auxiliary confirmation of the preliminary partial disappearance is partial disappearance, avoiding misjudgment caused by individual node detection errors. This step realizes the accurate judgment of the complete appearance and partial disappearance status of people in video image frames by introducing a combination of posture estimation model and semantic segmentation model, which has significant technical advantages. First of all, compared with traditional fixed frame rate sampling, this method starts with the complete entry of people into the frame and ends with partial exit, which can effectively avoid redundant interference of irrelevant frames and ensure the integrity and high correlation of the behavioral expression of the extracted frame sequence.Secondly, by judging the confidence and distribution range of key points, combined with the design of key node importance weights, the judgment of preliminary complete appearance and preliminary partial disappearance is made more structural and semantic, effectively improving the ability to capture key action moments. At the same time, a semantic segmentation mask auxiliary confirmation mechanism is further introduced to correct scenes where the posture estimation model may misjudge under conditions such as occlusion and blur, thereby enhancing the stability and robustness of the judgment. In addition, this method takes into account information in two dimensions: static structure (key points) and image area (mask pixels), realizing multi-feature collaborative verification, improving the accuracy of detection of the start and end of human behavior, and providing a high-quality data foundation for subsequent behavior feature extraction and trajectory construction. It has good practicality and engineering adaptability.

[0039] Next, considering the problems of insufficient recognition accuracy in low-light environments and single behavioral features in the prior art, the present invention adopts image enhancement and static and dynamic feature fusion solutions to perform image preprocessing and feature extraction, performs preprocessing operations on the image frame sequence, and extracts the behavioral features of human targets in the image frame sequence. The image frames are subjected to image enhancement processing, and the image enhancement processing includes image contrast enhancement, illumination normalization processing and edge denoising processing, which are used to improve the recognition accuracy of human areas in low-light environments; the human body key node information of each frame image is used to construct a static posture feature vector; further combined with the change trend of key points between adjacent image frames, the dynamic behavioral features of the human body are extracted to describe the action state of the human body in the frame image. Specifically, in the image preprocessing and feature extraction steps, the extracted image frame sequence is first subjected to image enhancement processing to adapt to the actual situation of uneven image quality in complex environments and to improve the stability and accuracy of subsequent recognition. This process involves three specific operations: First, image contrast enhancement is performed by using an adaptive histogram equalization algorithm to enhance details in dark areas and sharpen the edges of human figures. Second, illumination normalization is performed by using the Retinex algorithm to compensate for image illumination, addressing difficulties in human recognition caused by strong backlighting or partial overexposure. Third, edge-preserving noise reduction is performed by using bilateral filtering to smooth the image texture while preserving boundary information, facilitating subsequent object segmentation and keypoint detection. After image enhancement, the human pose feature construction phase begins. For each frame, a pre-trained human pose estimation model is used to identify multiple key nodes of the human body. The two-dimensional coordinates and confidence values ​​of these key nodes are obtained and encoded in a pre-set order into a static pose feature vector representing the structural pose of the human figure in the current frame. Furthermore, to incorporate temporal information, the keypoint positions of the current frame and several adjacent frames are differentiated or velocity-calculated to extract motion trends and action continuity features, forming a dynamic behavior feature vector that characterizes the motion state of the human figure in the current frame. This method comprehensively considers image quality enhancement and temporal behavior modeling, is suitable for actual monitoring environments such as low illumination, occlusion, and complex background, and has strong stability and practicality.

[0040] To address the issue of identity fragmentation in cross-camera scenarios and achieve continuous tracking of behavioral trajectories across the entire region, a multi-feature fusion cross-view matching scheme, namely, person trajectory reconstruction and identity association, is employed. The same person identified by different cameras is subjected to spatiotemporal identity matching to construct a cross-view person trajectory sequence. The specific operations include: utilizing human dynamic behavioral features, integrating appearance features with position information, and forming a joint feature representation for cross-camera identity matching. Based on the joint feature representation, identities of people appearing in different cameras are associated within multiple camera image frames to determine whether they belong to the same target individual. Upon successful identity association, a sequence of person trajectory sequences is constructed based on the temporal order, spatial location, and behavioral actions of the person appearing in different cameras. Specifically, for each detected human target in each image frame, dynamic behavioral features are extracted. These features consist of the displacement, angle change, and motion trajectory of key human posture points in adjacent image frames, describing the individual's short-term motion pattern. Simultaneously, static appearance features are extracted. A convolutional neural network is used to encode the human region image, capturing information such as color, texture, and clothing, representing the distinguishable appearance characteristics of the person. In addition, location information features are extracted, including camera number, image frame timestamp, and the target's relative position within the image, to provide spatiotemporal context. These three features are concatenated according to a unified structure to form a joint feature vector, which represents the state of the person appearing in a particular camera at each moment. Next, all person targets appearing in multiple camera image frame sequences are matched based on the joint feature representation. Cosine similarity or Euclidean distance is calculated, and temporal reachability and spatial proximity constraints are combined to determine whether they represent the same target individual. Specifically, when the joint feature similarity of the individuals identified by two cameras exceeds a set threshold, and the time difference between the two image frames is consistent with the physically reachable path between the two cameras, the system associates them as the same individual. To prevent mismatches, a bidirectional matching consistency mechanism is introduced, requiring both A and B to match A for a valid match to be considered. After identity association is completed, the target individuals are sorted according to their appearance time in each camera and, combined with their geographic location and action type information within each surveillance area, a continuous sequence of individual behavioral trajectories is constructed. The trajectory not only contains location and time, but also embeds behavioral status information, providing a semantically rich path data foundation for subsequent abnormal behavior analysis, with good temporal continuity and spatial traceability.

[0041] Finally, in order to solve the one-sidedness problem of traditional single action classification and achieve a deep understanding of behavioral intentions, the method of behavioral rule deviation analysis is used to calculate the abnormal behavior deviation score. The person's behavior trajectory is matched with the preset template, and the abnormal behavior deviation score is calculated. If the abnormal behavior deviation score exceeds the set deviation score threshold, the person is identified as an abnormal behavior person. The specific operations of matching the person's behavior trajectory with the preset template and calculating the abnormal behavior deviation score include: encoding the person's behavior trajectory into a behavior sequence sorted by time, and the behavior sequence is composed of position nodes, behavioral actions and behavior time; selecting the normal behavior template sequence that is closest to the person's behavior type and path structure, and calculating the trajectory matching score S based on time consistency, position path similarity and action sequence matching degree. match ; The calculation method of time consistency is Where N is the trajectory length, T t Represents the timestamp of time t, T t norm Represents the timestamp of template time t; the calculation method of location path similarity is Where M represents the total number of frames of the normal template, L t Represents the position feature at time t, represents the position feature of the template at time t; the degree of action sequence matching is calculated as follows: Where K represents the number of key action points, I() is the indicator function, which returns 1 if the condition is met and returns 0 if the condition is satisfied. Represents the behavioral action characteristics at time t, Represents the behavioral action features at template time t; the three are weighted fused to obtain the trajectory matching score. Multiple potential abnormal behavior segments are identified in the behavior trajectory, and each abnormal behavior segment is assigned an abnormal behavior confidence C j , combined with the duration to calculate the abnormal behavior confidence aggregation value S abn ; The confidence aggregate value S abn The calculation method is: Where J represents the number of potential abnormal behavior segments identified. Finally, the trajectory matching score is combined with the abnormal aggregation value to construct the final deviation score S deviation , calculated as: S deviation =(1-S match )·S abn Finally, a judgment operation is performed to determine if the abnormal behavior deviation score exceeds the set threshold, and the person is identified as an abnormal behavior person.

[0042] The above description is merely a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any person skilled in the art may utilize the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes for application in other fields. However, any simple modification, equivalent change, and modification of the above embodiments made in accordance with the technical essence of the present invention without departing from the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A method for identifying people with abnormal behavior by integrating multi-source surveillance images, characterized in that: The following steps are involved: S1. Video data acquisition: acquiring raw video data from multiple surveillance cameras deployed in the target surveillance area, wherein the video data includes image streams from different viewing angles and different time periods; S2. Video image frame extraction: performing inter-frame sampling on the original video stream, extracting an image frame sequence, and retaining the timestamp, camera number, and location information corresponding to each frame; S3, image preprocessing and feature extraction: preprocessing the image frame sequence and extracting the behavioral features of the human target in the image frame sequence; S4. Personnel trajectory reconstruction and identity association: The same person identified by different cameras is matched in time and space to construct a cross-viewpoint sequence of person behavior trajectories; S5. Behavioral rule deviation analysis: Match the person's behavior trajectory with the preset template and calculate the abnormal behavior deviation score. If the abnormal behavior deviation score exceeds the set deviation score threshold, the person is identified as an abnormal behavior person.

2. The method for identifying people with abnormal behavior by integrating multi-source surveillance images according to claim 1, characterized in that: The operation of performing inter-frame sampling on the original video stream in step S2 is to first determine that a person appears completely in the video stream as the start of inter-frame sampling, and to determine that a person disappears partially as the end of inter-frame sampling.

3. The method for identifying people with abnormal behavior by integrating multi-source surveillance images according to claim 2 is characterized in that: The method for judging complete appearance and partial disappearance is as follows: Key point network acquisition: Using posture estimation model to collect key nodes of the human body in real time; Preliminary complete appearance judgment: When all the key nodes collected are located in the valid area of ​​the image and the confidence of the key points is higher than the set confidence threshold, it is judged to be preliminarily complete appearance; Preliminary partial disappearance judgment: define the importance weights of key nodes and calculate the weighted ratio of disappeared nodes. When the disappearance ratio exceeds the set threshold, it is judged as preliminary partial disappearance; Enhanced judgment effect: Use the human body semantic segmentation model to segment the pixel area of ​​the human body and calculate the coverage rate of the human body mask and the image area. When the coverage rate is ≥90%, the auxiliary confirmation will judge the initial complete appearance as complete appearance; when the coverage rate is ≤40%, the auxiliary confirmation will judge the initial partial disappearance as partial disappearance.

4. The method for identifying people with abnormal behavior by integrating multi-source surveillance images according to claim 1, characterized in that: The image preprocessing and feature extraction operations in step S3 include: Performing image enhancement processing on the image frame, wherein the image enhancement processing includes image contrast enhancement, illumination normalization processing and edge denoising processing, so as to improve the recognition accuracy of human body areas in low-light environments; Using the key node information of the human body in each frame of image, a static posture feature vector is constructed; The dynamic behavior characteristics of the human body are further extracted by combining the change trend of key points between adjacent image frames to describe the action state of the human body in the frame image.

5. The method for identifying people with abnormal behavior by integrating multi-source surveillance images according to claim 1, characterized in that: The operations of reconstructing the personnel trajectory and associating it with the identity in step S4 include: Utilize the dynamic behavior characteristics of the human body, and introduce the appearance features and position information for fusion to form a joint feature representation for cross-camera identity matching; In multiple camera image frame sequences, the identities of people appearing in different cameras are associated based on joint feature representation to determine whether they belong to the same target individual; Based on the successful identity association, a sequence of people's behavior trajectories is constructed according to the time sequence, spatial position and behavioral actions of people appearing under different cameras.

6. The method for identifying people with abnormal behavior by integrating multi-source surveillance images according to claim 1, characterized in that: The specific operations of matching the personnel behavior trajectory with the preset template and calculating the abnormal behavior deviation score in step S5 include: Encode the behavior trajectory of the person into a behavior sequence sorted by time, wherein the behavior sequence consists of a location node, a behavior action, and a behavior time; Select the normal behavior template sequence that is closest to the person's behavior type and path structure, and calculate the trajectory matching score S based on time consistency, location path similarity and action sequence matching degree. match ; The calculation method of time consistency is Where N is the trajectory length, T t Represents the timestamp of time t, T t norm Represents the timestamp of template time t; the calculation method of location path similarity is Where M represents the total number of frames of the normal template, L t Represents the position feature at time t, represents the position feature of the template at time t; the degree of action sequence matching is calculated as follows: Where K represents the number of key action points, I() is the indicator function, which returns 1 if the condition is met and returns 0 if the condition is satisfied. Represents the behavioral action characteristics at time t, represents the action feature of the template at time t; the trajectory matching score is obtained by weighted fusion of the three; Identify multiple potential abnormal behavior segments in the behavior trajectory and assign an abnormal behavior confidence C to each abnormal behavior segment. j , combined with the duration to calculate the abnormal behavior confidence aggregation value S abn ; The confidence aggregate value S abn The calculation method is: Where J represents the number of identified potential abnormal behavior segments; Finally, the trajectory matching score and the anomaly aggregation value are combined to construct the final deviation score S deviation , calculated as: S deviation =(1-S match )·S abn .

Citation Information

Patent Citations

  • Abnormal behavior monitoring processing method and device, computer device and storage medium

    CN110443109A

  • Image processing method and system for intelligent security and protection monitoring

    CN118887622A

  • Financial escort process abnormal behavior identification method and system

    CN119559701A

  • Visualizing and updating learned trajectories in video surveillance systems

    US20110044498A1

Cited By

  • Multi-source data fusion-based unstable GPS data motion state identification method

    CN121256453A

  • Campus security monitoring method based on image recognition technology

    CN121486534A

  • Video monitoring abnormal behavior identification and tracking linkage method based on artificial intelligence

    CN121564045A