Abnormal behavior analysis method based on spatial positioning prior and multi-modal information fusion

By introducing spatial positioning priors and multimodal information fusion into abnormal behavior analysis technology, the problem of sensitive changes in single information sources and environmental changes in the existing technology is solved, and more accurate, robust and efficient abnormal behavior analysis is achieved.

CN120071209APending Publication Date: 2025-05-30GOSUNCN TECH GRP +1

Patent Information

Application Number
CN202510006576.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing anomaly behavior analysis techniques have the limitations of a single source of information when dealing with complex scenarios, making it difficult to fully capture visual, motion and spatial characteristics, and are sensitive to environmental changes, and are inefficient in processing long-term series data.

Method used

An abnormal behavior analysis method based on spatial positioning prior and multimodal information fusion is adopted. By obtaining image frames, human body skeleton information and spatial positioning information in the video stream, image information encoding and skeleton feature encoding are generated, and fusion and analysis are carried out through an abnormal behavior analysis network, and abnormal behavior warning information is finally generated.

Benefits of technology

It improves the accuracy and robustness of abnormal behavior analysis, can capture features in complex scenarios more comprehensively, enhance the ability to adapt to environmental changes, improve the efficiency of processing long-time series data, and realize real-time or near-real-time abnormal behavior analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071209A_ABST
    Figure CN120071209A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, computer vision and mode recognition, in particular to an abnormal behavior analysis method based on spatial positioning prior and multi-modal information fusion, which comprises the following steps: acquiring an image frame in a video stream; obtaining human body skeleton information in the image frame; acquiring spatial positioning information of a human body in the image frame; generating an image information code based on the image frame; generating skeleton feature codes based on the human skeleton information and the spatial positioning information; generating a fusion sequence based on the image information code and the skeleton feature code; generating an abnormal behavior analysis result through an abnormal behavior analysis network based on the fusion sequence; according to the method, abnormal behavior early warning information is generated based on an abnormal behavior analysis result, and by fusing image information, skeleton information and spatial positioning information, visual, motion and spatial features in a scene can be comprehensively captured, so that the ability of understanding complex behaviors is greatly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence, computer vision, and pattern recognition, and particularly relates to an abnormal behavior analysis method based on spatial positioning prior and multimodal information fusion. Background Art

[0002] With the development of society and the progress of technology, the demand for abnormal behavior analysis technology in the field of public security is increasing day by day. Abnormal behavior analysis has important practical significance and application value in the fields of security, traffic management, public place monitoring, etc. In recent years, this technical field has received extensive attention from the academic and industrial communities and has made remarkable progress.

[0003] Currently, abnormal behavior recognition technologies mainly include several methods. The method based on skeleton point analysis detects abnormal behaviors by extracting human skeleton points in images and analyzing the skeleton point trajectories. This method can effectively capture the details of human movement but often ignores the environmental context information. The methods based on convolutional neural networks (CNNs) and recurrent neural networks (RNNs) detect abnormal behaviors by extracting high-level features of the appearance and movement of pedestrians in videos. Such methods perform well in dealing with complex scenes but may be affected by occlusion and illumination changes. In recent years, abnormal behavior detection methods based on graph neural networks (GNNs) have begun to receive attention. This method uses the relationships between graph-structured data to model between objects or segments to complete the detection task and can better capture the interaction relationships between entities.

[0004] However, although these methods have achieved certain results in specific scenarios, they still face many challenges in practical applications. First, most existing methods rely on a single information source, which limits their ability to understand complex scenes. For example, methods that rely only on skeleton information may ignore important environmental clues, while methods based only on images may not be able to accurately capture the subtle changes in human movement. Second, existing methods generally have the problem of being sensitive to environmental changes. When the illumination conditions, camera angles, or scene layouts change, the performance of the system often drops significantly. In addition, many methods are inefficient in processing long-term sequence data and are difficult to meet the requirements of real-time monitoring.

[0005] The closest prior art is a Chinese invention patent with the publication number CN 115273040A and the title of "A Video-based Driving Behavior Analysis Method, Electronic Device, and Storage Medium". This method uses skeleton information as input and adopts a graph convolutional network to extract skeleton features. However, this method still has some obvious limitations. First, it only utilizes single skeleton information and ignores rich image semantic information and spatial location information, which may lead to incomplete and inaccurate analysis in complex scenarios. Second, this method uses spatio-temporal graph convolution and transformer graph convolution for feature extraction. Although it can capture certain spatio-temporal relationships, the effect may not be ideal when dealing with long time series and complex spatial relationships. Finally, this method is mainly aimed at driving behavior analysis and may be difficult to be directly applied to a wider range of abnormal behavior analysis scenarios.

[0006] In view of the deficiencies of the prior art, there is an urgent need for an abnormal behavior analysis method that can comprehensively utilize multi-modal information and consider spatial positioning prior at the same time. This method should be able to effectively fuse image, skeleton, and spatial location information to improve the accuracy and robustness of abnormal behavior analysis. At the same time, it should also be able to efficiently process long time series data and adapt to various complex monitoring scenarios. Summary of the Invention

[0007] The abnormal behavior analysis method based on the fusion of spatial positioning prior and multi-modal information proposed by the present invention is designed to solve the above technical problems. By innovatively fusing multiple information sources, introducing spatial positioning prior, and adopting advanced deep learning techniques, this method effectively improves the accuracy, robustness, and efficiency of abnormal behavior analysis.

[0008] The present invention protects an abnormal behavior analysis method based on the fusion of spatial positioning prior and multi-modal information, including:

[0009] An acquisition step, including:

[0010] Obtaining an image frame in a video stream;

[0011] Obtaining human skeleton information in the image frame;

[0012] Obtaining the spatial positioning information of the human body in the image frame;

[0013] A processing step, including:

[0014] Generating an image information encoding based on the image frame;

[0015] Generating a skeleton feature encoding based on the human skeleton information and the spatial positioning information;

[0016] Generating a fusion sequence based on the image information encoding and the skeleton feature encoding;

[0017] Based on the fusion sequence, generate an abnormal behavior analysis result through an abnormal behavior analysis network;

[0018] An output step, including:

[0019] Based on the abnormal behavior analysis result, generate an abnormal behavior warning message.

[0020] Preferably, the step of obtaining human skeleton information specifically includes:

[0021] Perform human detection on the image frame to obtain a human detection result;

[0022] Based on the human detection result, perform skeleton detection to obtain human skeleton information including joint point position information and joint point connection information.

[0023] Preferably, the step of obtaining spatial positioning information specifically includes:

[0024] Perform identity determination on the human body in the image frame to obtain identity information;

[0025] Based on the human skeleton information, estimate the landing point of the human body in three-dimensional space to obtain the landing point coordinates.

[0026] Preferably, the step of generating the image information encoding specifically includes:

[0027] Divide the image frame into multiple image blocks of a fixed size;

[0028] Perform linear mapping on each of the image blocks to obtain feature vectors;

[0029] Generate position encoding information;

[0030] Add the feature vectors and the position encoding information to obtain the image information encoding.

[0031] Preferably, the step of generating the skeleton feature encoding specifically includes:

[0032] Convert the human skeleton information into a tensor form;

[0033] Use a graph convolutional network to process the tensor to obtain a skeleton feature vector;

[0034] Fuse the skeleton feature vector and the spatial positioning information to obtain the skeleton feature encoding.

[0035] Preferably, the abnormal behavior analysis network is a network based on the Transformer architecture, including:

[0036] An encoder with a multi-head attention mechanism;

[0037] The decoder of the multi-head attention mechanism;

[0038] The multi-layer feed-forward neural network;

[0039] The layer normalization module.

[0040] Preferably, the multi-head attention mechanism specifically includes:

[0041] Mapping the input sequence to a query matrix, a key matrix, and a value matrix;

[0042] Calculating the attention weights;

[0043] Performing weighted summation on the value matrix based on the attention weights to obtain the attention output.

[0044] Preferably, it further includes a post-processing step:

[0045] Obtaining the identity information and the spatial positioning information in the forward processing module;

[0046] Converting the spatial positioning information to the coordinate position in the image coordinate system;

[0047] Matching the coordinate position with the human body position information in the abnormal behavior analysis result;

[0048] Assigning the identity information to the abnormal behavior analysis result that matches it.

[0049] Preferably, the output of the abnormal behavior analysis network includes:

[0050] The detected number of human bodies;

[0051] The position coordinates of each human body;

[0052] The abnormal behavior category corresponding to each human body.

[0053] Preferably, it further includes a model training step:

[0054] Obtaining a training data set with abnormal behavior annotations;

[0055] Based on the training data set, training the abnormal behavior analysis network by using the cross-entropy loss function;

[0056] Using the residual connection technology to combine the input features, the features after self-attention processing, and the features after perceptron processing to prevent the problems of gradient disappearance and gradient explosion.

[0057] The beneficial effects of the present invention are:

[0058] (1) By integrating image information, skeleton information, and spatial positioning information, this method can comprehensively capture visual, motion, and spatial features in the scene, greatly enhancing the ability to understand complex behaviors. This multi-modal fusion strategy effectively overcomes the limitations of a single information source, enabling the system to more accurately identify and analyze various abnormal behaviors.

[0059] (2) The present invention introduces a spatial positioning prior, providing important context information for behavior analysis. This not only improves the accuracy of analysis but also enhances the system's adaptability to environmental changes. For example, when monitoring different areas, the system can dynamically adjust the behavior judgment criteria according to the spatial position information, thereby achieving more intelligent and flexible detection of abnormal behaviors.

[0060] (3) The present invention adopts a Transformer-based network structure, which performs excellently in processing long sequence data. Through the self-attention mechanism, the system can effectively capture long-distance dependencies, thus better understanding and analyzing complex behavior patterns with longer durations. This feature gives this method a significant advantage in processing long-term surveillance videos and can timely detect potential abnormal behaviors.

[0061] (4) The method of the present invention exhibits excellent generalization ability and robustness in practical applications. By comprehensively utilizing multiple information sources and advanced deep learning technologies, this method can adapt to various complex surveillance scenarios, including conditions such as lighting changes and perspective changes. This greatly expands the application scope of the system, enabling it to play an important role in various public places, transportation hubs, important facilities, and other scenarios.

[0062] (5) The method of the present invention also has a significant improvement in computational efficiency. Through ingenious network design and optimization strategies, this method can achieve real-time or near-real-time analysis of abnormal behaviors, meeting the requirements of actual surveillance systems. This not only improves the practicality of the system but also makes large-scale deployment possible.

[0063] In summary, the method for abnormal behavior analysis based on spatial positioning prior and multi-modal information fusion proposed by the present invention innovatively solves the problems existing in the prior art, significantly improving the performance and practicality of abnormal behavior analysis. This method provides a powerful and flexible technical solution for the field of public safety and is expected to play an important role in enhancing social security levels and improving public management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 is the schematic diagram of the method of the present invention.

[0065] Figure 2 is the flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0066] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0067] As Figure 1-2 shown, the present invention provides an abnormal behavior analysis method based on spatial positioning prior and multi-modal information fusion. This method improves the accuracy and efficiency of abnormal behavior analysis by fusing multiple information sources. The specific implementation manners of this method will be described in detail below.

[0068] First of all, this method includes an acquisition step, a processing step and an output step. In the acquisition step, the method of the present invention acquires image frames in the video stream, human skeleton information in the image frames, and spatial positioning information of the human body. These information provide basic data support for subsequent abnormal behavior analysis.

[0069] Specifically, when acquiring image frames in the video stream, a high-frame-rate imaging device can be used, such as a camera with 30 frames per second or 60 frames per second, to ensure that sufficient detailed motion information is captured. Preferably, in an embodiment of the present invention, a high-definition camera with a resolution of 1920x1080 or higher can be used to obtain clear image details.

[0070] In the process of acquiring human skeleton information, this method first performs human detection on the image frames. Human detection can adopt deep learning algorithms, such as YOLO (You Only Look Once) or Faster R-CNN, etc. These algorithms can quickly and accurately locate the human body position in the image. After obtaining the human detection result, this method further performs skeleton detection to obtain human skeleton information including joint point position information and joint point connection information. Skeleton detection can use open-source tools such as OpenPose or AlphaPose, and these tools can effectively extract the key point information of the human body.

[0071] It should be noted that when performing human detection, in order to improve the detection accuracy, a confidence threshold can be set, such as 0.75. When the confidence of the detection result is higher than this threshold, it is considered that an effective human body is detected. The selection of this threshold is based on a large amount of experimental data, which can effectively reduce the false detection rate while ensuring the detection accuracy.

[0072] Next, this method obtains the spatial positioning information of the human body. This step includes determining the identity of the human body in the image frame to obtain identity information, and estimating the landing point of the human body in the three-dimensional space based on the human skeleton information to obtain the landing point coordinates. Identity determination can be achieved using face recognition technology or gait recognition technology. For example, a deep learning model such as FaceNet can be used for face feature extraction and matching to achieve identity recognition.

[0073] When estimating the landing point coordinates, this method uses the human skeleton information, especially the position information of the foot joint points, combines the camera parameters and the scene geometry information, and calculates the actual position of the human body in the three-dimensional space through techniques such as inverse perspective transformation. Preferably, in an embodiment of the present invention, a binocular camera or a depth camera can be used to obtain more accurate depth information, thereby improving the estimation accuracy of the landing point coordinates.

[0074] The acquisition of spatial positioning information is crucial for abnormal behavior analysis. It can not only help the system understand the position and movement trajectory of the human body in space, but also provide important context information for subsequent behavior analysis. For example, in a monitoring scenario, if a person stays in a specific area for a long time, it may need attention.

[0075] Through the above acquisition steps, this method successfully obtains multi-modal information, including visual information (image frames), motion information (skeleton information), and spatial information (positioning information). These rich information sources provide comprehensive data support for subsequent abnormal behavior analysis, helping to improve the accuracy and robustness of the analysis.

[0076] In the next processing step, this method first generates an image information encoding based on the image frame. The purpose of this step is to convert the high-dimensional image data into a feature representation suitable for processing by machine learning models. The specific encoding process will be described in detail later.

[0077] Then, this method generates a skeleton feature encoding based on the human skeleton information and the spatial positioning information. This step organically combines the posture information and the spatial position information of the human body to form a richer and more meaningful feature representation.

[0078] Next, this method fuses the image information encoding and the skeleton feature encoding to generate a fusion sequence. This step is one of the core innovations of this method. It realizes the effective integration of multi-modal information, enabling subsequent abnormal behavior analysis to utilize visual, motion, and spatial information simultaneously.

[0079] Finally, this method generates an abnormal behavior analysis result based on the fusion sequence through an abnormal behavior analysis network. The specific structure and working principle of this network will be described in detail later.

[0080] In the output step, the method generates abnormal behavior warning information based on the abnormal behavior analysis results. These warning information may include important information such as the type of abnormal behavior, the time and location of occurrence, and the identities of the persons involved, providing timely and effective decision-making support for security management personnel.

[0081] Through the above steps, the method of the present invention realizes efficient and accurate analysis of abnormal behaviors. The advantages of this method lie in the full utilization of multi-modal information, especially the introduction of spatial location priors, which greatly improves the accuracy and robustness of the analysis. At the same time, by adopting advanced deep learning technologies, this method can handle complex scenarios and behavior patterns and has strong practical value.

[0082] In the method of the present invention, generating the image information encoding is a key step, which converts high-dimensional image data into a feature representation suitable for processing by a deep learning model. Specifically, this step includes several important sub-processes.

[0083] First, the method divides the image frame into multiple image patches of a fixed size. This segmentation strategy is called "patch embedding" and is inspired by the Vision Transformer model. In a preferred embodiment of the present invention, the size of each image patch is set to 16x16 pixels. This size selection is based on a trade-off between computational efficiency and feature expression ability. Smaller image patches can capture more detailed local features but will increase computational complexity; larger image patches may lose some detailed information. For example, an image of size HxWxC is first divided into N non-overlapping patches of size 16*16*3 according to a fixed size.

[0084] Next, a linear mapping is performed on each image patch to obtain a feature vector. This step is usually implemented through a fully connected layer, which maps each image patch to a vector space of a fixed dimension. Preferably, in the method of the present invention, the dimension of this vector space can be set to 256 or 512. A higher dimension can retain more information but will also increase the computational amount of subsequent processing. The dimension of each patch block is flattened into a one-dimensional vector, and then these vectors are linearly mapped E into a space of a fixed dimension D. Therefore, after passing through the linear projection layer, tokens of size N*D are obtained.

[0085] Then, the method generates position encoding information. The purpose of position encoding is to provide each image patch with its position information in the original image, because after the image is segmented into patches, the original spatial relationship information is lost. In one embodiment of the present invention, a sine position encoding method can be adopted. The advantage of this encoding method is that it can process sequences of arbitrary length, and the encoding result has good interpolation properties. Since the picture order information is missing, position encoding is needed. The position encoding can be obtained through sine input or training embedding, and the size of the position encoding is 1*N.

[0086] Finally, the feature vector is added to the position encoding information to obtain the final image information encoding. This addition operation allows the model to consider both the content and position information when processing each image patch; for example, adding the tokens and the position encoding to obtain an embedding of size N*D.

[0087] In the process of generating the skeleton feature encoding, the method of the present invention first converts the human skeleton information into a tensor form. The purpose of this step is to convert the skeleton data into a format suitable for processing by a deep learning model. In a preferred embodiment, the two-dimensional coordinates and confidence information of each joint point can be combined into a three-dimensional vector, and then the information of all joint points is concatenated into a tensor.

[0088] Next, the method uses a graph convolutional network (GCN) to process this tensor to obtain the skeleton feature vector. The use of GCN is an important innovation point of the present invention, which can effectively capture the spatial relationship in the skeleton structure. Obtain the human skeleton information in the picture, where the skeleton information includes the position information of each joint point in the picture and the connection information of each joint point. Then, the above information is represented in the form of a tensor and sent into the graph convolution in the spatial dimension, and then a residual connection is made with the input. After batch normalization, the feature vector V is obtained. In one embodiment of the present invention, two or three layers of GCN can be used, and the output dimension of each layer can be set to 64 or 128. This setting can extract rich enough skeleton features while maintaining computational efficiency.

[0089] Then the feature vector V is flattened into a one-dimensional vector and these vectors are linearly mapped E into a space of a fixed dimension. Therefore, after passing through the linear projection layer, a token of size n*D is obtained.

[0090] Next, the three-dimensional spatial coordinate points of the global coordinates, that is, the three-dimensional spatial position information of the human grounding point, are position-encoded. The position encoding can be obtained through sine input or training embedding, and finally an n-dimensional feature matrix is obtained. Finally, the tokens and the position encoding are added to obtain an embedding of size n*D.

[0091] This method fuses the skeleton feature vector with the spatial location information to obtain the final skeleton feature encoding. This fusion process can be achieved through a simple concatenation operation or by using a more complex attention mechanism. Preferably, an adaptive fusion strategy can be adopted in the method of the present invention to dynamically adjust the weights of the skeleton features and the spatial location information according to different scenarios.

[0092] In the method of the present invention, the abnormal behavior analysis network adopts a network structure based on the Transformer architecture. This choice is based on the excellent performance of the Transformer in sequence modeling tasks. The core advantage of the Transformer architecture lies in its ability to effectively capture long-distance dependencies, which is crucial for understanding complex human behavior sequences.

[0093] Specifically, the abnormal behavior analysis network of the present invention includes an encoder and a decoder with a multi-head attention mechanism, a multi-layer feed-forward neural network, and a layer normalization module. In a preferred embodiment, the encoder and the decoder can each include 6 Transformer layers, and the number of hidden units in each layer can be set to 512. This setting can achieve a good balance between model capacity and computational efficiency.

[0094] The multi-head attention mechanism is the core component of the Transformer architecture. In the method of the present invention, the implementation of the multi-head attention mechanism includes several key steps. First, the input sequence is mapped into a query matrix, a key matrix, and a value matrix. This mapping process is usually achieved through a linear transformation. In an embodiment of the present invention, 8 attention heads can be used, and the dimension of each head can be set to 64. This setting allows the model to learn information from different representation subspaces, thereby enhancing the feature extraction ability.

[0095] Then, the attention weights are calculated based on the query matrix and the key matrix. This calculation process usually adopts the dot-product attention mechanism, that is, multiplying the query matrix by the transpose of the key matrix and then normalizing it through the softmax function. Preferably, a temperature parameter can be introduced in the method of the present invention to adjust the smoothness of the softmax, for example, setting the temperature parameter to sqrt(d_k), where d_k is the dimension of the key vector.

[0096] Finally, a weighted sum of the value matrix is performed based on the attention weights to obtain the attention output. This step realizes the aggregation of information according to the attention distribution. In the embodiment of the present invention, a linear transformation layer can be added after the attention output to increase the expressive ability of the model.

[0097] Specifically, the image embedding and the skeleton embedding are concatenated to obtain a sequence of size (N + n) * D, and then the sequence is fed into the transformer network model. The transformer network model includes a transformer encoder and a transformer decoder that apply the multi-head attention mechanism, as well as a multi-layer feed-forward neural network and layer normalization.

[0098] The multi-head attention mechanism first maps the input matrix Z through WSA to obtain Q, K, and V. Then Q, K, and V are divided into i parts. The multi-head self-attention mechanism can be expressed as:

[0099] MSA(Q, K, V) = Concat(head 1 ,..., head i ) * W 0

[0100] where concat means concatenating the feature tensors. headi represents the i-th single-head attention head.

[0101] The multi-head attention consists of i single-head attention mechanisms. Among them, the single-head self-attention mechanism is expressed as:

[0102]

[0103] where Q, K, and V are obtained from the input sequence X through a series of matrix multiplication transformations, representing the query, key, and value respectively. The Softmax function is used to calculate the attention weights аi,j, and the final result is the sum of the values weighted by the attention weights.

[0104] In this embodiment, we use an 8-head self-attention module to capture multi-scale features. To avoid bias in the activation layer due to excessive dimensions, each self-attention module undergoes appropriate scale conversion. The outputs of these modules are aggregated to form a feature matrix, which is then processed by regularization and input into the perceptron layer. During this process, this embodiment adopts the residual connection technology to combine the features input to the Transformer, the features processed by the multi-head self-attention, and the features processed by the multi-layer perceptron, thereby preventing the problems of gradient disappearance and gradient explosion.

[0105] After a series of processing by the Transformer encoder and decoder, we obtain the output sequence. This sequence is then fed into the feed-forward neural network, and finally, the position information with a dimension of Nobj × 2 and the corresponding abnormal behavior categories are output. Here, Nobj represents the number of detected human bodies, and 2 represents the coordinate dimensions of each human body position.

[0106] Through the steps described in detail above, the method of the present invention can effectively extract and fuse multi-modal information, providing a powerful feature representation for abnormal behavior analysis. This Transformer-based network structure can not only capture complex spatio-temporal dependencies but also has good scalability, being able to adapt to abnormal behavior analysis tasks of different scales. In the method of the present invention, the post-processing step is crucial for generating the final early warning information of abnormal behavior. This step cleverly correlates the identity information and spatial location information obtained in the forward processing module with the abnormal behavior analysis results, thereby providing a more accurate and meaningful warning output.

[0107] Specifically, the method first obtains the identity information and spatial location information in the forward processing module. These information are obtained in the early stage of video stream analysis and contain the identity identifiers and initial spatial positions of each human body in the scene. It should be noted that in practical applications, the acquisition of identity information may be restricted by privacy protection regulations. Therefore, in a preferred embodiment of the present invention, anonymization processing can be adopted, such as using randomly generated unique identifiers to replace the real identity information.

[0108] Next, the method converts the spatial location information into coordinate positions in the image coordinate system. The purpose of this step is to map the position information in three-dimensional space onto a two-dimensional image plane so as to match the human body position information in the abnormal behavior analysis results. In the embodiment of the present invention, a perspective transformation matrix can be used to achieve this conversion. Preferably, the camera calibration parameters can be utilized to construct an accurate perspective transformation matrix to ensure the accuracy of the conversion.

[0109] Then, the method matches the converted coordinate positions with the human body position information in the abnormal behavior analysis results. This matching process usually adopts the nearest neighbor algorithm, that is, finding the nearest converted coordinate position for the human body position in each abnormal behavior analysis result. In a preferred embodiment of the present invention, a distance threshold can be set, such as 5% of the image width. Only when the matching distance is less than this threshold is the match considered successful. This can effectively avoid false matches and improve the robustness of the system.

[0110] Specifically, first extract the identity information and spatial location information corresponding to the human body from the forward processing module. Then convert the spatial location information into coordinate positions in the image coordinate system, and then match this position information with the human body position information output by the model. After successful matching, assign the identity information to the corresponding human body output by the model. Finally, the post-processing module outputs the early warning of abnormal behavior with identity information for each human body in the figure.

[0111] Finally, the present method assigns identity information to the matching abnormal behavior analysis results. Through this step, the system can generate abnormal behavior warnings containing specific identity information, greatly improving the practicality and operability of the warning information.

[0112] In terms of the output of the abnormal behavior analysis network, the method of the present invention designs a comprehensive and effective output structure. Specifically, the network output includes the number of detected human bodies, the position coordinates of each human body, and the abnormal behavior category corresponding to each human body.

[0113] The number of detected human bodies provides important reference information for subsequent processing. In an embodiment of the present invention, a maximum detection number can be set, such as 10 or 20, to balance the computational efficiency and the scene complexity. When the actual number of detected human bodies exceeds this threshold, the system can preferentially process the detection results with higher confidence.

[0114] The position coordinates of each human body are usually represented in the image coordinate system, that is, the (x, y) coordinate pair. Preferably, in the method of the present invention, these coordinates can be normalized to the range of [0, 1] to facilitate conversion and comparison between images with different resolutions.

[0115] For the abnormal behavior category corresponding to each human body, the present method uses the multi-label classification method for output. This means that a person may exhibit multiple abnormal behaviors simultaneously. In a preferred embodiment of the present invention, a series of abnormal behavior categories can be predefined, such as breaking into a restricted area, loitering, fighting, etc., and a confidence score is output for each category. By setting an appropriate confidence threshold, such as 0.7, the system can screen out high-likelihood abnormal behaviors for alarm.

[0116] Finally, the method of the present invention also includes an important model training step, which is crucial for ensuring the performance of the system. First, the present method obtains a training data set with abnormal behavior annotations. The quality and scale of this data set directly affect the performance of the model. In an embodiment of the present invention, data from multiple sources can be used, including real surveillance videos and simulated abnormal behavior data, to increase the diversity and representativeness of the data.

[0117] Next, based on the training data set, the present method trains the abnormal behavior analysis network using the cross-entropy loss function. The cross-entropy loss function performs well in multi-classification problems and can effectively guide the model to learn to distinguish different categories of abnormal behaviors. Preferably, the class-weighted cross-entropy loss can be used to balance the problem of uneven distribution of different abnormal behavior categories in the data set.

[0118] During the training process, the method of the present invention uses the residual connection technique to combine the input features, the features processed by self-attention, and the features processed by the perceptron. This technique effectively alleviates the problems of gradient disappearance and gradient explosion in deep networks, enabling the model to train deeper network structures more stably. In a preferred embodiment of the present invention, a residual connection can be added after each Transformer block and used in conjunction with layer normalization to further improve the training stability and the generalization ability of the model.

[0119] Through the steps and techniques described in detail above, the abnormal behavior analysis method based on spatial positioning prior and multi-modal information fusion provided by the present invention can effectively identify and warn of abnormal behaviors in various complex scenarios. This method not only has innovation in technical implementation, but also shows excellent performance and reliability in practical applications, providing a powerful and flexible solution for the field of public safety.

[0120] It is easy for those skilled in the art to understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An abnormal behavior analysis method based on spatial positioning prior and multimodal information fusion, characterized in that: include: The acquisition steps include: Get the image frame in the video stream; Acquire human skeleton information in the image frame; Acquiring spatial positioning information of a human body in the image frame; Processing steps include: Based on the image frame, generate image information code; Generate skeleton feature code based on the human skeleton information and the spatial positioning information; Generate a fusion sequence based on the image information encoding and the skeleton feature encoding; Based on the fusion sequence, an abnormal behavior analysis result is generated through an abnormal behavior analysis network; the output step includes: Based on the abnormal behavior analysis result, abnormal behavior warning information is generated.

2. The method according to claim 1, characterized in that The step of obtaining human skeleton information specifically includes: Performing human body detection on the image frame to obtain a human body detection result; Based on the human body detection result, skeleton detection is performed to obtain human body skeleton information including joint point position information and joint point connection information.

3. The method according to claim 1, characterized in that The step of obtaining spatial positioning information specifically includes: Performing identity determination on a human body in the image frame to obtain identity information; Based on the human skeleton information, the landing point of the human body in the three-dimensional space is estimated to obtain the coordinates of the landing point.

4. The method according to claim 1, characterized in that: The step of generating image information coding specifically includes: Dividing the image frame into a plurality of image blocks of fixed sizes; Performing linear mapping on each of the image blocks to obtain a feature vector; generating position encoding information; The feature vector is added to the position coding information to obtain image information coding.

5. The method according to claim 1, characterized in that The step of generating skeleton feature coding specifically includes: Converting the human skeleton information into a tensor form; Processing the tensor using a graph convolutional network to obtain a skeleton feature vector; The skeleton feature vector is fused with the spatial positioning information to obtain a skeleton feature code.

6. The method according to claim 1, characterized in that The abnormal behavior analysis network is a network based on the Transformer architecture, including: Multi-head attention encoder; Decoder with multi-head attention mechanism; Multi-layer feed-forward neural network; Layer Normalization Module.

7. The method according to claim 6, characterized in that The multi-head attention mechanism specifically includes: Map the input sequence into a query matrix, a key matrix, and a value matrix; Calculate attention weights; The value matrix is ​​weighted summed based on the attention weights to obtain the attention output.

8. The method according to claim 1, characterized in that It also includes a post-processing step: obtaining identity information and spatial positioning information in the forward processing module; Convert the spatial positioning information to a coordinate position in an image coordinate system; Matching the coordinate position with the human body position information in the abnormal behavior analysis result; The identity information is assigned to the abnormal behavior analysis result that matches it.

9. The method according to claim 1, characterized in that: The outputs of the abnormal behavior analysis network include: The number of human bodies detected; The position coordinates of each human body; Abnormal behavior category corresponding to each human body.

10. The method according to claim 1, characterized in that It also includes the model training steps: obtaining a training dataset with abnormal behavior annotations; Based on the training data set, the abnormal behavior analysis network is trained using a cross entropy loss function; The residual connection technology is used to combine the input features, the features after self-attention processing, and the features processed by the perceptron to prevent the gradient disappearance and gradient explosion problems.

Citation Information

Patent Citations

  • Video-based driving behavior analysis method, electronic equipment and storage medium

    CN115273040A

Cited By

  • Behavior recognition and abnormity early warning method, device and system applied to terrace classroom

    CN121661408A

  • Behavior identification and abnormity early warning method and system

    CN121768084A

  • A behavior recognition and abnormality early warning method and system

    CN121768084B