An abnormal behavior recognition method, device, equipment and medium
Patent Information
- Application Number
- CN202610867559.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]具体而言,单个采集设备的视角狭窄,难以覆盖目标的完整行为轨迹
[0025]第二方面至第四方面中任意一种实现方式所带来的技术效果可参见第一方面的实现方式所带来的技术效果,此处不再赘述。
Smart Images

Figure CN122821458A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device and medium for identifying abnormal behavior. Background Technology
[0002] In open intelligent surveillance systems (such as plazas, transportation hubs, and smart parks), the key challenge is how to automatically and accurately identify abnormal behavior from video streams from multiple acquisition devices (such as cameras).
[0003] Specifically, a single acquisition device has a narrow field of view, making it difficult to cover the complete behavioral trajectory of the target. Furthermore, the analysis of multiple acquisition devices often remains at the level of simple image stitching or fusion after independent analysis, failing to achieve semantic-level information collaboration and global correlation reasoning across acquisition devices.
[0004] Therefore, the accuracy of identifying collaborative abnormal behavior across acquisition devices is low and the generalization ability is poor. Summary of the Invention
[0005] To address the problems in the prior art, embodiments of this application provide an abnormal behavior recognition method, apparatus, device, and medium to improve the accuracy and generalization ability of abnormal behavior recognition in open scenarios.
[0006] In a first aspect, embodiments of this application provide an abnormal behavior identification method, the method comprising: Initial features of a target individual are extracted from video streams from multiple acquisition devices, and a heterogeneous spatiotemporal graph is constructed based on the initial features. The heterogeneous spatiotemporal graph includes: target nodes representing the target individual, region nodes representing the acquisition area, device nodes representing the acquisition devices, and connecting edges used to characterize the relationships between nodes. The connecting edges are classified into temporal type and spatial type according to the type of relationship they represent. For each target node in the heterogeneous spatiotemporal graph, the following operations are performed using the heterogeneous graph neural network model: at each moment within the observation time window, based on the weights and types of the connecting edges corresponding to the target node, the node features of the target node and the node features of its neighboring nodes are weighted and fused to obtain the behavioral state features of the target node at that moment. The behavioral state features of the target node at each time step are aggregated to obtain a comprehensive feature; wherein, the comprehensive feature is used to characterize the overall behavioral pattern of the target individual within the observation time window; Based on the deviation between the comprehensive features and the preset normal behavior features, abnormal behavior identification results are generated for the target individual.
[0007] This application constructs a heterogeneous spatiotemporal graph containing target nodes, region nodes, and device nodes by classifying connection edges into temporal and spatial types. This connects anonymous individuals scattered across different acquisition devices with their dynamic scenes, overcoming the limitations of single-device perspectives and the information silos of multiple devices. Based on the weights and types of connection edges, the target node's own features and those of its neighboring nodes are weighted and fused, so that the generated behavioral state features incorporate both spatial and temporal association information. By aggregating the behavioral state features at each time step into a comprehensive feature and calculating the deviation from preset normal behavioral features, anomaly identification is transformed into a distance measurement problem in the feature space. This mechanism can capture the temporal evolution details of behavior and determine whether behavior deviates from normal behavior as a whole, achieving global perception and semantic-level understanding across device scenarios, significantly improving the accuracy and generalization ability of abnormal behavior identification in open scenarios.
[0008] In one possible implementation, for each target node, the node features of the target node include: appearance features of the corresponding target individual, position features for characterizing the spatial position of the corresponding target individual, and motion features for characterizing the motion state of the corresponding target individual; For each regional node, the node characteristics of the regional node include: dynamic environmental characteristics, which include the density, movement speed and movement direction of the target individuals within the corresponding acquisition area; For each device node, the node characteristics include: the installation location of the corresponding acquisition device, the intrinsic parameter matrix, and the rotation matrix and translation vector used to characterize the orientation of the device.
[0009] By setting node features for target nodes, region nodes, and device nodes respectively, the data mapping path from the original video stream to the heterogeneous spatiotemporal graph is clearly defined. Specifically, the appearance features of target nodes are used for cross-device re-identification, position features for spatial proximity judgment and boundary establishment, and motion features for motion trajectory entropy calculation and abnormal behavior detection. The dynamic environmental features of region nodes enable real-time updates of scene states and their participation in inference. The position features and intrinsic parameter matrices of device nodes are used to convert image coordinates to world coordinates, supporting accurate spatial distance calculation and device field-of-view determination. This multi-dimensional node feature setting provides a rich heterogeneous information foundation for subsequent weighted fusion based on the weight and type of connecting edges.
[0010] In one possible implementation, the connecting edge includes a spatial edge belonging to a first spatial type, a temporal edge belonging to a time type, an observation edge belonging to a second spatial type, and a belonging edge belonging to a third spatial type; The spatial edge is used to connect different target nodes observed by the same acquisition device at the same time, and the distance between the target individuals they represent meets the preset conditions. The time edge is used to connect different target nodes representing the same target individual observed by different acquisition devices at different times; The observation edge is used to connect a device node to the target node corresponding to the target individual observed by the acquisition device represented by the device node; The belonging edge is used to connect a target node to the region node corresponding to the current collection area of the target individual represented by the target node.
[0011] By establishing four types of connection edges with different semantics—spatial edges, temporal edges, observation edges, and attribution edges—a heterogeneous graph structure capable of comprehensively depicting four core relationships in a monitoring scenario was constructed. Specifically, spatial edges carry local interaction information between spatially adjacent targets at the same time, temporal edges carry trajectory association information of the same target across devices and time periods, observation edges carry device observation quality information, and attribution edges carry interaction information between the target and the scene. This refined edge classification lays the foundation for the subsequent spatiotemporal decoupling message propagation mechanism.
[0012] In one possible implementation, the weight of each belonging edge is determined as follows: Based on the target individual represented by the target node connected by the home edge, the weight of the home edge is determined by the dwell time and motion trajectory entropy within the collection area represented by the regional node connected by the home edge; wherein, the motion trajectory entropy is used to characterize the degree of disorder of the motion pattern of the target individual within the collection area.
[0013] By determining dynamic weights for the attribution edges based on dwell time and motion trajectory entropy, precise quantification of the interaction intensity between targets and the scene is achieved. The longer a target stays within the collection area, or the higher its motion trajectory entropy (i.e., the more chaotic its movement pattern, exhibiting wandering or turning back), the greater the weight of the attribution edge. This dynamic weight, as a crucial component of the connection edge confidence, ensures that targets with high anomaly potential receive greater attention in subsequent weighted fusion, thereby enabling early focusing on subtle anomalous behaviors such as wandering and loitering.
[0014] In one possible implementation, the step of weightedly fusing the node features of the target node and the node features of its neighboring nodes based on the weights and types of the connecting edges corresponding to the target node includes: Based on the type of each connecting edge corresponding to the target node, each neighbor node is divided into the following types of neighbor nodes: each first spatial neighbor node connected to the target node through a spatial edge, each temporal neighbor node connected to the target node through a temporal edge, each second spatial neighbor node connected to the target node through an observation edge, and each third spatial neighbor node connected to the target node through a belonging edge. Based on the weights of each connecting edge corresponding to the target node, the attention coefficients of each type of neighbor node are determined respectively. Based on the determined attention coefficients, the node features of the corresponding types of neighboring nodes are weighted and aggregated to obtain the first spatial feature, the temporal feature, the second spatial feature, and the third spatial feature. The first spatial feature, the temporal feature, the second spatial feature, the third spatial feature, and the node feature of the target node are fused together.
[0015] By categorizing neighbor nodes into first-space neighbors (other target nodes), temporal neighbors (same target nodes across devices), second-space neighbors (device nodes), and third-space neighbors (region nodes), and determining attention coefficients based on the weights of the connecting edges before weighted aggregation, a message propagation mechanism achieving spatiotemporal decoupling and scene fusion is implemented. First-space neighbors carry local interaction information, temporal neighbors carry cross-device trajectory information, second-space neighbors carry observation quality information, and third-space neighbors carry scene state information. The behavioral state features generated by fusing these four types of features with the target's own features can comprehensively reflect the behavioral state of the target individual and its spatiotemporal scene context.
[0016] In one possible implementation, after obtaining the behavioral state features of the target node at that moment, and before aggregating the behavioral state features of the target node at each moment to obtain a comprehensive feature, the method further includes: The behavioral state features at this moment are fused with the behavioral state features generated by the target node at historical moments prior to this moment; The fusion result is used as the behavioral state feature of the target node after the update at that moment.
[0017] By configuring a memory enhancement operation before temporal aggregation, the behavioral state features at the current moment are fused with features from historical moments, ensuring that the updated behavioral state features at each moment contain the target's own historical behavioral information. This operation enables the model to capture long-range temporal dependencies in behavior, such as understanding the complete behavioral evolution chain from "entering the area" to "brief stay" and then to "long-term wandering," thus allowing subsequent temporal aggregation operations to be performed based on features that already contain historical information.
[0018] In one possible implementation, the aggregation of the behavioral state features of the target node at each time step to obtain comprehensive features includes: Based on the correlation between the behavioral state features of the target node at each time step, the importance weight of the behavioral state features of the target node at each time step is determined. Based on the determined importance weights, the behavioral state features of the target node at the corresponding time are weighted and summed to obtain the comprehensive features.
[0019] By performing temporal aggregation of behavioral state features at each time step, determining importance weights based on the correlation between features at each time step, and then performing weighted summation, the model can automatically focus on the most critical moments in the behavioral sequence for anomaly detection. Moments with significant changes in behavioral patterns or exhibiting anomalous characteristics receive higher weights, while moments with stable behavioral patterns and low information content receive correspondingly lower weights, thereby generating more discriminative comprehensive features.
[0020] In one possible implementation, the preset normal behavior characteristics are different, and different regional nodes correspond to different normal behavior characteristics; The step of generating an identification result for the abnormal behavior of the target individual based on the deviation between the comprehensive features and preset normal behavior features includes: Based on the current region node to which the target individual belongs, the normal behavior feature corresponding to the current region node is called from multiple normal behavior features corresponding to different region nodes, and recorded as the target normal behavior feature; Based on the deviation between the comprehensive features and the target's normal behavioral features, an abnormal behavior identification result is generated for the target individual.
[0021] By maintaining corresponding normal behavior features for different regional nodes and calculating deviation based on the target's current region during the identification phase, scene-adaptive anomaly identification is achieved. Since different collection areas (such as entrances / exits, gathering areas, and passageways) have different definitions of "normal behavior," using a unified normal behavior benchmark is difficult to adapt to changing monitoring scenarios. By independently constructing corresponding normal behavior features for each regional node, the system can dynamically switch judgment benchmarks based on the target's current specific region, thereby significantly improving the accuracy and generalization ability of anomaly identification in different scenarios.
[0022] Secondly, embodiments of this application provide an abnormal behavior identification device, the device comprising: A construction unit is used to extract initial features of a target individual from video streams from multiple acquisition devices and construct a heterogeneous spatiotemporal graph based on the initial features; wherein, the heterogeneous spatiotemporal graph includes: target nodes representing the target individual, region nodes representing the acquisition area, device nodes representing the acquisition devices, and connecting edges used to characterize the relationships between nodes; the connecting edges are divided into temporal type and spatial type according to the type of relationship they characterize. The feature fusion unit is used to perform the following operations for each target node in the heterogeneous spatiotemporal graph through a heterogeneous graph neural network model: at each time point within the observation time window, based on the weights and types of each connecting edge corresponding to the target node, the node features of the target node and the node features of each neighboring node of the target node are weighted and fused to obtain the behavioral state features of the target node at that time point. The feature aggregation unit is used to aggregate the behavioral state features of the target node at each time step to obtain a comprehensive feature; wherein the comprehensive feature is used to characterize the overall behavioral pattern of the target individual within the observation time window; The behavior recognition unit is used to generate abnormal behavior recognition results for the target individual based on the deviation between the comprehensive features and preset normal behavior features.
[0023] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and when the computer program is executed by the processor, it implements the method described in any one of the abnormal behavior recognition methods of the first aspect.
[0024] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method described in any one of the abnormal behavior identification methods of the first aspect.
[0025] The technical effects of any of the implementation methods in the second to fourth aspects can be found in the technical effects of the implementation method in the first aspect, and will not be repeated here. Attached Figure Description
[0026] Figure 1 A flowchart illustrating an abnormal behavior identification method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of an abnormal behavior recognition device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0028] The terms "first" and "second" in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising" and any variations thereof are intended to cover non-exclusive protection. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. The term "multiple" in this application can mean at least two, for example, two, three, or more, and the embodiments of this application do not impose limitations.
[0029] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These embodiments should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that in the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solutions of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0030] The acquisition, transmission, storage, and use of data in this application all comply with the requirements of relevant national laws and regulations.
[0031] Before introducing the abnormal behavior identification method provided in the embodiments of this application, the technical background of the embodiments of this application will be described in detail below for ease of understanding.
[0032] In open intelligent surveillance systems (such as plazas, transportation hubs, and smart parks), the key challenge is how to automatically and accurately identify abnormal behavior from video streams from multiple acquisition devices (such as cameras).
[0033] Specifically, a single acquisition device has a narrow field of view, making it difficult to cover the complete behavioral trajectory of the target. Furthermore, the analysis of multiple acquisition devices often remains at the level of simple image stitching or fusion after independent analysis, failing to achieve semantic-level information collaboration and global correlation reasoning across acquisition devices.
[0034] Therefore, the accuracy of identifying collaborative abnormal behavior across acquisition devices is low and the generalization ability is poor.
[0035] In view of this, embodiments of this application provide an abnormal behavior recognition method, apparatus, device, and medium to improve the accuracy and generalization ability of abnormal behavior recognition in open scenarios.
[0036] The following is combined Figure 1 This application will introduce an abnormal behavior identification method according to an embodiment of the present application.
[0037] refer to Figure 1 , Figure 1 An exemplary method for identifying abnormal behavior is provided in an embodiment of this application. The method includes the following steps: Step S101: Extract initial features of the target individual from the video streams of multiple acquisition devices, and construct a heterogeneous spatiotemporal map based on the initial features.
[0038] The heterogeneous spatiotemporal graph includes: target nodes representing target individuals, region nodes representing the collection area, device nodes representing the collection equipment, and connecting edges used to characterize the relationships between nodes. The connecting edges are divided into temporal type and spatial type according to the type of relationship they represent.
[0039] (a) Extraction of initial features of the target individual: For each video stream acquired by a capture device, target detection algorithms (such as YOLO, Faster R-CNN, etc.) and tracking algorithms (such as DeepSORT, etc.) can be used to detect and track individual targets within it.
[0040] For each detected target individual, its initial features are extracted, including: extracting the appearance features of the corresponding target individual, for example, by extracting feature vectors through a person re-identification (ReID) network to distinguish different target individuals; extracting the spatial location features of the corresponding target individual, which can be the three-dimensional spatial coordinates of the target individual in the world coordinate system, obtained by transforming image coordinates using the intrinsic and extrinsic parameter matrices of the acquisition device; and extracting motion features to characterize the motion state of the corresponding target individual, which are the instantaneous velocity vectors of the target individual, calculated from the position difference between adjacent frames.
[0041] (II) Node construction of heterogeneous spacetime graphs: Based on the extracted initial features, a heterogeneous spatiotemporal graph is constructed. This heterogeneous spatiotemporal graph contains three types of heterogeneous nodes: The first category is target nodes, which represent target individuals (such as pedestrians, vehicles, etc.) in the monitoring scene. Each target individual corresponds to one target node. The node features of the target nodes are set according to the initial features extracted above, including appearance features. Location features and motion characteristics, such as instantaneous velocity. .
[0042] The second category is regional nodes, which represent the collection area in the monitoring scenario (such as entrances / exits, gathering areas, passages, and other areas with specific semantic meanings). The node characteristics of regional nodes include static geographic information (such as regional boundary coordinates and regional type codes) and dynamic environmental characteristics. Among them, the dynamic environmental characteristics are updated in real time and include: the density of target individuals within the collection area (calculated by the number of target nodes falling into the area), movement speed, and movement direction.
[0043] The third category is device nodes, which represent physical acquisition devices (such as cameras). The node characteristics of device nodes include: the installation location of the acquisition device (coordinates in the world coordinate system), the intrinsic parameter matrix (focal length, principal point coordinates, etc.), and the rotation matrix and translation vector used to characterize the orientation of the device.
[0044] (III) Construction of connecting edges in heterogeneous spacetime graphs: This heterogeneous spatiotemporal graph also includes connecting edges used to represent the relationships between nodes. These connecting edges are divided into temporal and spatial types according to the type of relationship they represent.
[0045] Among them, time-type connection edges are used to connect nodes with temporal correlation (such as different target nodes representing the same target individual observed by different acquisition devices at different times), and spatial connection edges are used to connect nodes with spatial correlation (such as different target nodes observed by the same acquisition device at the same time, and the distance between the target individuals they represent meets the preset conditions).
[0046] Specifically, the connecting edges include the following four types: The first type is spatial edges, belonging to the first spatial type. Spatial edges are used to connect different target nodes observed by the same acquisition device at the same time, whose distances between the target individuals they represent satisfy preset conditions. If two target nodes... and Satisfying the following conditions: 1) belonging to the same acquisition device view; 2) having the same timestamp; 3) spatial distance. Then, an undirected spatial edge is established. The weight of the spatial edge decays as the distance increases, and its initial weight... Calculated using the Gaussian decay formula:
[0047] in, It is the spatial distance between the target individuals represented by the two target nodes. It is the distance decay factor.
[0048] The second type is the time edge, which belongs to the time type. A time edge is used to connect different target nodes representing the same target individual observed by different acquisition devices at different times. For candidate node pairs ( Calculate the association confidence level using the following steps. : First, calculate multi-source evidence. Appearance similarity. for and Cosine similarity of appearance features; confidence level of motion continuity Calculation based on the continuity of motion trajectory; spatiotemporal consistency Calculated using Gaussian form:
[0049] in, It is the time difference between the appearance of the two target nodes. The prior passage time for the scene. This represents the time standard deviation.
[0050] Then, the appearance similarity, motion continuity confidence, and spatiotemporal consistency are concatenated into [[ Input a lightweight adaptive fusion network for adaptive fusion:
[0051] in, , , , For learnable parameters, σ is the ReLU activation function, and σ is the Sigmoid function.
[0052] Finally, if (Based on a preset similarity threshold), a directed time edge is then established. Its weight .
[0053] The third type is the observation edge, which belongs to the second space type. An observation edge is used to connect a device node to the target node corresponding to the target individual observed by the acquisition device represented by that device node. The target is determined by the intrinsic and extrinsic parameters of the acquisition device. Is its three-dimensional spatial position in Within the field of view cone. If so, then establish from... point to The directed edges. Their initial weights. The centrality of the target individual in the imaging plane is calculated to characterize the observation quality.
[0054] The fourth type is the attribution edge, which belongs to the third space type. Attribution edges are used to connect a target node to the region node corresponding to the current collection area of the target individual represented by that target node. If... fall into Within the defined acquisition area, a collection point is established from... point to The directed edges. Their weights. Dynamic initialization is performed based on the intensity of the interaction or the potential for anomalies.
[0055] in, It is the current dwell time of the target individual represented by the target node in this area. It represents the trajectory entropy of the target individual within the collection area. The longer the dwell time or the more chaotic the trajectory (the higher the trajectory entropy), the greater the weight of the belonging edge.
[0056] (iv) Differences in the connection edges of the same target node at different times: It should be noted that for the same target node, the type and weight of its connecting edges will change dynamically at different times. The following example illustrates this: Suppose that the target individual P moves in a square monitoring scene. The collection area includes the "entrance and exit" (referred to as collection area A) and the "gathering area" (referred to as collection area B). The collection devices include cameras Cam1 and Cam2.
[0057] At the first moment, P is within the field of view of Cam1, located in acquisition area A, with other target individuals Q nearby. At this time, a spatial edge exists between the target node corresponding to P and the target node corresponding to Q; due to their proximity, this spatial edge has a high weight. Since P has no historical trajectory information, there is no temporal edge. An observation edge exists between the target node corresponding to P and the Cam1 node; because P is located near the center of the image, the observation quality is good, and the weight of this observation edge is 0.85. A belonging edge exists between the target node corresponding to P and the area node representing acquisition area A; because P has just entered acquisition area A and its dwell time is short, the weight of this belonging edge is low.
[0058] At the second moment, P leaves Cam1's field of view and enters the blind zone between Cam1 and Cam2. At this time, the target node corresponding to P has no spatial edge (because there are no other target nodes nearby), and since P is not within the field of view of any acquisition device, there is no observation edge. Simultaneously, since P is not within any acquisition area, there is also no belonging edge.
[0059] At the third moment, P enters Cam2's field of view, is located within acquisition area B, and begins to wander. At this time, the target node corresponding to P has no spatial edge (no other target nodes nearby). However, there is a temporal edge between P and the target node from the first moment, with a weight of 0.89 (calculated based on appearance similarity and motion continuity). There is an observation edge between P and the Cam2 node; since P is located at the edge of the image, the observation quality is generally low, and the weight of this observation edge is 0.70. There is a belonging edge between P and the regional node representing acquisition area B; as P's dwell time within acquisition area B increases and its motion trajectory entropy rises, the weight of this belonging edge gradually increases.
[0060] At the fourth time point, P is still within the acquisition area B, continuously lingering. At this time, the target node corresponding to P still has no spatial edge. The temporal edge connects to the node at the first time point and also to the node at the third time point (the continuous trajectory of the same target). The observation edge remains connected to the Cam2 node, with a weight of 0.70. The attribution edge connects to the region node representing acquisition area B; because P stays in region B for a long time and its movement trajectory is chaotic (lingering), the weight of the attribution edge reaches 0.92.
[0061] This shows that the types and number of neighboring nodes connected to the same target node, as well as the weights of each connecting edge, change dynamically at different times, reflecting the dynamic characteristics of heterogeneous spatiotemporal graphs.
[0062] Step S102: For each target node in the heterogeneous spatiotemporal graph, the following operation is performed using the heterogeneous graph neural network model: At each moment within the observation time window, based on the weights and types of the connecting edges corresponding to the target node, the node features of the target node and the node features of its neighboring nodes are weighted and fused to obtain the behavioral state features of the target node at that moment.
[0063] (a) Classification of neighboring nodes: Based on the type of connecting edges, the neighboring nodes of the target node are divided into the following four categories: The first category is the first spatial neighbor nodes, which are other target nodes connected through spatial edges. These neighbor nodes represent other spatially adjacent target individuals using the same acquisition device at the same time, and carry local interaction context information.
[0064] The second category is time-neighbor nodes, which are other target nodes connected via time edges. These neighbor nodes represent other nodes representing the same target individual at different acquisition devices and at different times, carrying cross-acquisition device context information.
[0065] The third category is second-space neighbor nodes, which are device nodes connected through observation edges. These neighbor nodes represent the acquisition devices that observed the target individual and carry observation quality context information.
[0066] The fourth type is third-space neighbor nodes, which are region nodes connected by belonging edges. These neighbor nodes represent the current collection area of the target individual and carry scene state context information.
[0067] (ii) Attention mechanism of confidence perception: This application employs a confidence-aware heterogeneous attention mechanism for weighted fusion. For nodes... and its neighboring nodes The connection edge type is First, calculate a confidence-modulated attention coefficient. :
[0068] in, They are nodes The current feature vector; It is for edge type The dedicated linear transformation weight matrix; It is for edge type Attention vector; This indicates vector concatenation; It is a cross-flow confidence gating factor used to inject prior confidence based on the specific semantics of the edge.
[0069] The methods for determining the confidence gating factors for various types of edges are as follows: For time edge ( ), confidence gating factor It can be calculated using the following formula:
[0070] in, It is the cosine similarity of the appearance features of the target individuals represented by two target nodes; It is the normalized time difference between the occurrence of the target individuals represented by the two target nodes; It is the cosine of the angle between the optical axes of the two acquisition devices; It is a learnable parameter vector.
[0071] For the observed edge ( ), confidence gating factor The value is directly taken as the initial weight of the observed edge. .
[0072] For the spatial edge ( ) and belonging edge ( ), its confidence gating factor Directly associate or set as the corresponding initial edge weight .
[0073] (III) Layered processing of weighted fusion: This application divides the weighted fusion operation into two major branches: time-weighted fusion and spatial-weighted fusion, and then further refines spatial-weighted fusion into three subclasses.
[0074] Temporal weighted fusion is performed on temporal neighbor nodes connected by temporal edges. Specifically, attention coefficients for each temporal neighbor node are determined based on the weights of the temporal edges. Then, based on the determined attention coefficients, the node features of each temporal neighbor node are weighted and aggregated to obtain the temporal features.
[0075] Spatial weighted fusion is performed on spatial neighbor nodes connected by spatial type edges, and is further subdivided into three subclasses: The first subclass targets the first spatial neighbor nodes (other target nodes connected by spatial edges). Based on the weight of the spatial edge (i.e., the spatial distance between two target nodes), the attention coefficients for each first spatial neighbor node are determined, and then weighted aggregation is performed to obtain the first spatial features.
[0076] The second subclass targets the second spatial neighbor nodes (device nodes connected via observation edges). Based on the weights of the observation edges, attention coefficients for each second spatial neighbor node are determined, and then weighted aggregation is performed to obtain the second spatial features.
[0077] The third subclass targets third-space neighbor nodes (region nodes connected by belonging edges). Based on the weight of the belonging edges, the attention coefficients for each third-space neighbor node are determined, and then weighted aggregation is performed to obtain the third-space features.
[0078] (iv) Feature fusion: Finally, the temporal features obtained by time-weighted fusion and the three sub-features (first spatial feature, second spatial feature, and third spatial feature) obtained by spatial-weighted fusion are fused with the node features of the target node itself to obtain the behavioral state features of the target node at that moment.
[0079] (v) Memory enhancement operations: After obtaining the behavioral state features of the target node at that moment, a memory augmentation operation can be optionally performed. After each round of graph convolution and feature update, a memory unit is introduced for each target node. This unit fuses the node's current features with its historical behavioral state memory to generate an updated memory.
[0080] in, It is the memory of the historical behavior and state of target node i at the previous moment. It is a characteristic of the current moment. It is the updated memory, serving as the final feature representation of the node at the current moment. This operation enables the behavioral state features at each moment to include the target's own historical behavioral information, thereby possessing the ability to capture long-term temporal dependencies of behavior.
[0081] Step S103: Aggregate the behavioral state features of the target node at each time step to obtain comprehensive features.
[0082] Among them, the comprehensive features are used to characterize the overall behavioral patterns of the target individual within the observation time window.
[0083] This application employs a self-attention mechanism for temporal aggregation. Within the time window T, the behavioral state memory sequence of the target node i output by the memory enhancement module is { }, aggregated into a fixed-length comprehensive feature .
[0084] First, the attention score at each time step is calculated using learnable parameters:
[0085] Then, the attention score is normalized using the Softmax function to obtain the importance weight at each time step:
[0086] Finally, the behavioral state features at each time step are weighted and summed according to their importance to obtain the comprehensive features:
[0087] Where W, b, and v are learnable model parameters. This represents the importance weight of the memory at time t. This comprehensive feature is essentially a compact "behavioral fingerprint" of the target individual's behavioral patterns throughout the entire observation time window.
[0088] For example, continuing the previous example, individual P wanders within area B. Assume the observation window is 10 seconds. During the first 8 seconds, P walks normally within area B, with similar behavioral characteristics at each moment. In the last 2 seconds, P begins to wander back and forth, exhibiting a significant change in behavioral characteristics. After calculating the correlation between characteristics at each moment using a self-attention mechanism, it is found that the characteristics of the first 8 seconds are highly correlated, representing a stable pattern of normal behavior; the characteristics of the last 2 seconds have a lower correlation with the characteristics of the first 8 seconds, representing a sudden pattern of abnormal behavior. Therefore, the characteristics of the last 2 seconds receive a higher weight, while the characteristics of the first 8 seconds have a lower weight. The weighted summation of the comprehensive characteristics will be primarily dominated by the abnormal behavioral characteristics of the last 2 seconds, thus more accurately representing P's "wandering" behavior pattern.
[0089] Step S104: Based on the deviation between the comprehensive features and the preset normal behavior features, generate abnormal behavior identification results for the target individual.
[0090] (a) Normal behavioral characteristics of scene adaptation: There are multiple preset normal behavior characteristics, and different normal behavior characteristics correspond to different regional nodes. This is because different data collection areas have different definitions of "normal behavior". For example, the normal behavior characteristics of the "entrance / exit" area describe the pattern of "quick straight passage". A short stay at the entrance / exit may be normal, but a long stay may be abnormal. On the other hand, the normal behavior characteristics of the "gathering area" area describe the pattern of "normal stay and slow movement". A long stay in the gathering area may be normal, but quick passage may be abnormal.
[0091] During the identification phase, based on the current region node to which the target individual belongs, the normal behavior feature corresponding to the current region node is retrieved from multiple normal behavior features corresponding to different region nodes and recorded as the target normal behavior feature.
[0092] (ii) Pre-construction of normal behavioral characteristics: The pre-defined normal behavioral characteristics are constructed in advance in the following way: Multiple individuals exhibiting normal behavior are identified as normal samples. For each normal sample, the following steps are performed: initial feature extraction, construction of a heterogeneous spatiotemporal graph, generation of behavioral state features at each time step using a heterogeneous graph neural network model, and aggregation of these behavioral state features into a comprehensive feature set. This yields the comprehensive feature set for each normal sample. Next, the multi-source credibility information of the sample node corresponding to each normal sample in the heterogeneous spatiotemporal graph is determined. This multi-source credibility information includes the weights of the connecting edges corresponding to the sample node (particularly the weights of the observation and belonging edges) and the dynamic environmental features of the nodes in the region to which the sample node belongs. Then, the multi-source credibility information is input into a pre-defined lightweight network to generate the contribution weights of the normal sample. Based on the contribution weights of each normal sample, the comprehensive feature set of each normal sample is weighted and aggregated, and the result of this weighted aggregation is used as the normal behavior feature set.
[0093] Specifically, for K normal samples, the contribution weight of each sample... Calculated using a learnable weight assignment network:
[0094] in, The comprehensive features of sample node k, The observation confidence level for aggregation. Let k be the feature vector of the region to which sample node k belongs. For the Sigmoid function, It is the ReLU activation function. , , These are learnable parameters.
[0095] Then, based on the contribution weight of each normal sample, the comprehensive features of each normal sample are weighted and aggregated to obtain the normal behavior features:
[0096] This pre-construction mechanism allows high-quality, high-scene-representative normal samples to dominate the construction of normal behavior benchmarks, while the weights of marginal or low-quality samples are suppressed, thus obtaining a purer and more discriminative normal behavior benchmark.
[0097] (III) Deviation Calculation and Recognition Result Generation: The deviation between the comprehensive features and the target's normal behavioral features is calculated. This application uses cosine distance as the measure of deviation:
[0098] in, The comprehensive characteristics of the target individual to be tested. The target's normal behavioral characteristics are to be invoked. Represents the dot product. This represents the L2 norm. :when When this occurs, it indicates that the behavior is completely consistent with the normal pattern; The larger the value, the greater the degree of deviation and the greater the possibility of an anomaly.
[0099] The calculated deviation is compared with a preset threshold: if the deviation is greater than the threshold, the target individual is judged to have abnormal behavior; otherwise, it is judged to be normal. The final output of the abnormal behavior identification result may include the target individual's identifier, abnormal score, abnormal type (such as "long-term loitering"), associated region and other information.
[0100] For example, target individual P wanders within collection area B (cluster area), and its comprehensive characteristics represent a behavioral pattern of "slow speed and reversing trajectory." The normal behavioral characteristics corresponding to collection area B are then retrieved, which represent a pattern of "normal stillness and slow movement" (i.e., a straight trajectory rather than reversing). The deviation between the two is calculated. Since there is a significant difference between "reversing trajectory" and "slow movement," the deviation is relatively high, for example, 0.685. Assuming a preset threshold of 0.6, since 0.685 is greater than 0.6, target individual P is determined to exhibit abnormal behavior, specifically "wandering," and the corresponding identification result is output.
[0101] Based on the foregoing description, in this embodiment, by dividing the connection edges into temporal and spatial types, a heterogeneous spatiotemporal graph containing target nodes, region nodes, and device nodes is constructed, linking anonymous individuals scattered across different acquisition devices and their dynamic scenes, overcoming the problems of limited single-device perspective and information silos among multiple devices; by dividing weighted fusion into temporal weighted fusion and spatial weighted fusion, and further subdividing spatial weighted fusion into first spatial features (target-target interaction), second spatial features (device observation quality), and third spatial features (scene state), a comprehensive modeling of the spatiotemporal context of the target individual is achieved; by based on the connection edges... The system weights and types of features are used to weight and fuse the target node's own features with the features of four types of neighboring nodes. This allows the generated behavioral state features to incorporate both temporal correlation information (cross-device trajectory history) and multi-dimensional spatial correlation information (local interaction, observation quality, scene state). Memory enhancement operations are used to fuse current-moment features with historical-moment features, ensuring that each moment's features contain the target's own historical behavioral information, enabling the capture of long-term temporal dependencies in behavior. A self-attention mechanism aggregates behavioral state features from each moment into comprehensive features and calculates the deviation from preset normal behavioral features, transforming anomaly identification into a distance measurement problem in the feature space. This mechanism captures the temporal evolution details of behavior and judges whether behavior deviates from normal behavior as a whole, achieving global perception and semantic-level understanding in cross-device scenarios. It significantly improves the accuracy and generalization ability of anomaly behavior identification in open scenarios (especially anomalies requiring temporal judgment and scene understanding, such as loitering, wandering, and cross-device collaborative anomalies). Furthermore, it requires only a small number of normal samples for deployment, greatly reducing the system's construction, maintenance, and expansion costs.
[0102] Based on the same inventive concept, this application provides an abnormal behavior recognition device. Please refer to... Figure 2 The device includes: a construction unit 201, a feature fusion unit 202, a feature aggregation unit 203, and a behavior recognition unit 204, wherein: The construction unit 201 is used to extract initial features of the target individual from the video streams of multiple acquisition devices and construct a heterogeneous spatiotemporal graph based on the initial features. The heterogeneous spatiotemporal graph includes: target nodes representing the target individual, region nodes representing the acquisition area, device nodes representing the acquisition devices, and connecting edges used to characterize the relationship between nodes. The connecting edges are divided into temporal type and spatial type according to the type of relationship they represent.
[0103] The feature fusion unit 202 is used to perform the following operations for each target node in the heterogeneous spatiotemporal graph through the heterogeneous graph neural network model: at each time point within the observation time window, the node features of the target node and the node features of each neighbor node of the target node are weighted and fused according to the weight of each connecting edge corresponding to the target node and the type of each connecting edge, so as to obtain the behavioral state features of the target node at that time point.
[0104] The feature aggregation unit 203 is used to aggregate the behavioral state features of the target node at each time step to obtain comprehensive features; wherein, the comprehensive features are used to characterize the overall behavioral pattern of the target individual within the observation time window.
[0105] The behavior recognition unit 204 is used to generate abnormal behavior recognition results for a target individual based on the deviation between the comprehensive features and the preset normal behavior features.
[0106] The abnormal behavior recognition device provided in this application embodiment and the abnormal behavior recognition method in the above embodiment have the same beneficial effects, and will not be described in detail here.
[0107] Having introduced the abnormal behavior recognition method and abnormal behavior recognition device according to exemplary embodiments of this application, the electronic device provided according to embodiments of this application will now be described.
[0108] This application provides an electronic device that can implement the abnormal behavior recognition method described above. Please refer to... Figure 3 The device includes a memory 301, a processor 302, and a bus 303.
[0109] The memory 301 is used to store computer programs executed by the processor 302. The memory 301 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0110] Memory 301 may be volatile memory, such as random-access memory (RAM); memory 301 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 301 may be any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 301 may be a combination of the above-mentioned memories.
[0111] The processor 302 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 302 is used to implement the abnormal behavior recognition method in the above embodiments when it invokes a computer program stored in the memory 301.
[0112] This application embodiment does not limit the specific connection medium between the memory 301 and the processor 302 described above. This application embodiment... Figure 3 The memory 301 and the processor 302 are connected via a bus 303, and the bus 303 is in Figure 3 The connections between other components are shown in thick lines only and are not intended to be limiting. Bus 303 can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0113] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium. The computer program product includes computer program code, which, when executed on a computer, causes the computer to perform any of the abnormal behavior identification methods discussed above. Since the principle by which the above-described computer-readable storage medium solves the problem is similar to that of the abnormal behavior identification method, the implementation of the above-described computer-readable storage medium can be referred to the implementation of the method, and repeated details will not be elaborated further.
[0114] Based on the same inventive concept, this application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to execute any of the abnormal behavior identification methods discussed above. Since the principle of the above computer program product in solving the problem is similar to that of the abnormal behavior identification method, the implementation of the above computer program product can refer to the implementation of the method, and repeated details will not be described again.
[0115] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0116] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0117] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0118] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of user-operated steps to be executed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0119] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for identifying abnormal behavior, characterized in that, The method includes: Initial features of a target individual are extracted from video streams from multiple acquisition devices, and a heterogeneous spatiotemporal graph is constructed based on the initial features. The heterogeneous spatiotemporal graph includes: target nodes representing the target individual, region nodes representing the acquisition area, device nodes representing the acquisition devices, and connecting edges used to characterize the relationships between nodes. The connecting edges are classified into temporal type and spatial type according to the type of relationship they represent. For each target node in the heterogeneous spatiotemporal graph, the following operations are performed using the heterogeneous graph neural network model: at each moment within the observation time window, based on the weights and types of the connecting edges corresponding to the target node, the node features of the target node and the node features of its neighboring nodes are weighted and fused to obtain the behavioral state features of the target node at that moment. The behavioral state features of the target node at each time step are aggregated to obtain a comprehensive feature; wherein, the comprehensive feature is used to characterize the overall behavioral pattern of the target individual within the observation time window; Based on the deviation between the comprehensive features and the preset normal behavior features, abnormal behavior identification results are generated for the target individual.
2. The method according to claim 1, characterized in that, For each target node, the node features of the target node include: the appearance features of the corresponding target individual, the position features used to characterize the spatial position of the corresponding target individual, and the motion features used to characterize the motion state of the corresponding target individual; For each regional node, the node characteristics of the regional node include: dynamic environmental characteristics, which include the density, movement speed and movement direction of the target individuals within the corresponding acquisition area; For each device node, the node characteristics include: the installation location of the corresponding acquisition device, the intrinsic parameter matrix, and the rotation matrix and translation vector used to characterize the orientation of the device.
3. The method according to claim 1, characterized in that, The connecting edges include spatial edges belonging to the first spatial type, temporal edges belonging to the time type, observation edges belonging to the second spatial type, and belonging edges belonging to the third spatial type. The spatial edge is used to connect different target nodes observed by the same acquisition device at the same time, and the distance between the target individuals they represent meets the preset conditions. The time edge is used to connect different target nodes representing the same target individual observed by different acquisition devices at different times; The observation edge is used to connect a device node to the target node corresponding to the target individual observed by the acquisition device represented by the device node; The belonging edge is used to connect a target node to the region node corresponding to the current collection area of the target individual represented by the target node.
4. The method according to claim 3, characterized in that, For each belonging edge, the weight of the belonging edge is determined as follows: Based on the target individual represented by the target node connected by the home edge, the weight of the home edge is determined by the dwell time and motion trajectory entropy within the collection area represented by the regional node connected by the home edge; wherein, the motion trajectory entropy is used to characterize the degree of disorder of the movement pattern of the target individual within the collection area.
5. The method according to claim 3 or 4, characterized in that, The step of weightedly fusing the node features of the target node and the node features of its neighboring nodes based on the weights and types of the connecting edges corresponding to the target node includes: Based on the type of each connecting edge corresponding to the target node, each neighbor node is divided into the following types of neighbor nodes: each first spatial neighbor node connected to the target node through a spatial edge, each temporal neighbor node connected to the target node through a temporal edge, each second spatial neighbor node connected to the target node through an observation edge, and each third spatial neighbor node connected to the target node through a belonging edge. Based on the weights of each connecting edge corresponding to the target node, the attention coefficients of each type of neighbor node are determined respectively. Based on the determined attention coefficients, the node features of the corresponding types of neighboring nodes are weighted and aggregated to obtain the first spatial feature, the temporal feature, the second spatial feature, and the third spatial feature. The first spatial feature, the temporal feature, the second spatial feature, the third spatial feature, and the node feature of the target node are fused together.
6. The method according to claim 1, characterized in that, After obtaining the behavioral state features of the target node at that moment, and before aggregating the behavioral state features of the target node at each moment to obtain comprehensive features, the method further includes: The behavioral state features at this moment are fused with the behavioral state features generated by the target node at historical moments prior to this moment; The fusion result is used as the behavioral state feature of the target node after the update at that moment.
7. The method according to claim 1, characterized in that, The process of aggregating the behavioral state features of the target node at each time step to obtain comprehensive features includes: Based on the correlation between the behavioral state features of the target node at each time step, the importance weight of the behavioral state features of the target node at each time step is determined. Based on the determined importance weights, the behavioral state features of the target node at the corresponding time are weighted and summed to obtain the comprehensive features.
8. The method according to claim 1, characterized in that, The preset normal behavior characteristics are different, and different regional nodes correspond to different normal behavior characteristics; The step of generating an identification result for the abnormal behavior of the target individual based on the deviation between the comprehensive features and preset normal behavior features includes: Based on the current region node to which the target individual belongs, the normal behavior feature corresponding to the current region node is called from multiple normal behavior features corresponding to different region nodes, and recorded as the target normal behavior feature; Based on the deviation between the comprehensive features and the target's normal behavioral features, an abnormal behavior identification result is generated for the target individual.
9. An abnormal behavior recognition device, characterized in that, The device includes: A construction unit is used to extract initial features of a target individual from video streams from multiple acquisition devices and construct a heterogeneous spatiotemporal graph based on the initial features; wherein, the heterogeneous spatiotemporal graph includes: target nodes representing the target individual, region nodes representing the acquisition area, device nodes representing the acquisition devices, and connecting edges used to characterize the relationships between nodes; the connecting edges are divided into temporal type and spatial type according to the type of relationship they characterize. The feature fusion unit is used to perform the following operations for each target node in the heterogeneous spatiotemporal graph through a heterogeneous graph neural network model: at each time point within the observation time window, based on the weights and types of each connecting edge corresponding to the target node, the node features of the target node and the node features of each neighboring node of the target node are weighted and fused to obtain the behavioral state features of the target node at that time point. The feature aggregation unit is used to aggregate the behavioral state features of the target node at each time step to obtain a comprehensive feature; wherein the comprehensive feature is used to characterize the overall behavioral pattern of the target individual within the observation time window; The behavior recognition unit is used to generate abnormal behavior recognition results for the target individual based on the deviation between the comprehensive features and preset normal behavior features.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the method of any one of claims 1 to 8.