Abnormal Behavior Supervision Method and System for Interrogation Scenarios Based on ST-Transformer Network

The human skeleton extraction and abnormal behavior detection are carried out through the ST-Transformer network, which solves the privacy and accuracy of abnormal behavior recognition in interrogation scenarios, and achieves efficient and accurate supervision of abnormal behaviors.

CN115953831BActive Publication Date: 2025-07-11SHANGHAI JIAOTONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211580460.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-07-11
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify abnormal behaviors in interrogation scenarios, and there are problems of private information leakage and low recognition accuracy.

Method used

Using the method based on the ST-Transformer network, a video database of interrogation scenes is constructed through human skeleton extraction and abnormal behavior detection, and an abnormal behavior analysis is performed using the ST-Transformer network model to protect privacy and improve identification accuracy.

Benefits of technology

It realizes abnormal behavior detection with high privacy and high recognition accuracy in interrogation scenarios, and can output and update detection results in real time, which is suitable for a variety of scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115953831B_ABST
    Figure CN115953831B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for supervising abnormal behaviors in an interrogation scenario based on an ST-Transformer network, including establishing a training database and a test database for interrogation scenario video samples; constructing an ST-Transformer network model, using the training database as input to train the model to obtain a trained ST-Transformer network model; acquiring and decoding the video stream to be detected to obtain the image of the current frame, determining whether there is a human body in the image, if not, continuing to acquire the next frame of the image; if so, extracting the human body skeleton position in the image and saving the human body skeleton sequence of the preset frames; sending it to the trained ST-Transformer network model for abnormal behavior detection and analysis; saving the results of the detection and analysis and the corresponding data to the test database for users to view and supervise. The present invention extracts the human body skeleton from the original video to generate a video containing only the human body skeleton, effectively avoiding noise interference. At the same time, the original video information in the surveillance video is not stored in the system, which has high privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of abnormal behavior detection, and in particular, to a method and system for supervising abnormal behavior in an interrogation scenario based on an ST-Transformer network. Background Art

[0002] In recent years, the number of monitors in cities has developed rapidly, and a large amount of monitored video data has been generated by monitoring systems deployed in various scenarios. In the monitored video, analyzing the behavior of the main body therein, that is, people, requires a large amount of manpower and becomes very difficult in the face of massive data. Therefore, relying on computer vision methods under artificial intelligence to intelligently analyze the monitored video and detect abnormal events therein, including key information such as fighting, crowd gathering, and falling, is of great significance for actual social stability and intelligent prevention and control.

[0003] Early abnormal behavior detection relied on manual pixel features in the video, such as optical flow feature extraction, motion trajectory capture, human body structure estimation, local binarization, etc. Such methods can have good effects on fixed data sets. However, their judgment depends on prior design, and for variable scenarios, individuals are prone to missed detection and false detection. With the development of artificial intelligence, deep learning technology has been more and more widely applied to the recognition of abnormal behaviors. By first training an end-to-end neural network, inputting the perceived image and video into a structure with layers of convolutions, and extracting high-level semantic information into feature maps with layers of deep expressions, the features of these abnormal behaviors can be represented by the vectors of the last layer, and finally a fully connected layer is used for classification and recognition. The basic networks commonly used for video data include 3D-CNN networks or RNN networks. Compared with traditional methods, deep learning networks do not require manual design of features, but learn feature expressions from the data through the network. Therefore, after training with a large amount of data, the model can well fit an intelligent detection model parameter. However, this method also obtains information from pixel points, so it is more sensitive to interference such as weather, environment, and background, and there are still limitations in the detection robustness of the model.

[0004] Patent document CN113076772A discloses a method for recognizing abnormal behaviors based on full modalities, including collecting full-modal behavior samples and constructing an abnormal behavior sample library; constructing abnormal behavior features through an algorithm, automatically identifying abnormal behavior personnel and tracking their action trajectories, and notifying relevant personnel at each site in real time to take corresponding countermeasures, such as timely discovering abnormal behaviors of people: sudden running, falling, chasing and beating, etc., and reminding the manager.

[0005] However, patent document CN113076772A integrates multi-modal detection methods, including voice, expression, etc. It cannot well protect the privacy information of relevant personnel, and the accuracy of abnormal behavior recognition is low.

[0006] Patent document CN113269103A discloses a method and system for abnormal behavior detection based on a spatial graph convolutional network, belonging to the technical field of machine vision processing. It extracts the skeleton feature space graph of all individuals in the video frame to be detected; uses the trained abnormal score model to process the spatial graph to obtain the abnormal score of each skeleton in the video frame; performs a max pooling operation on the abnormal scores of all skeletons in the video frame to obtain the abnormal score of the video frame; and classifies the abnormal behavior level of the video frame according to the high and low abnormal scores of the video frame.

[0007] However, the target scenario of patent document CN113269103A is different from that of the present invention, and this document is not suitable for application in scenarios with strong privacy, such as interrogation scenarios. At the same time, the present invention uses a different ST-Transformer network for recognition. This network can better capture the correlation between human joint points through the multi-head attention in the Transformer structure to distinguish abnormalities of the personnel in the interrogation room.

[0008] The PGT network was proposed in the paper "Pyramid spat ial-temporal graph transformer for skeleton-basedaction recognition", and it uses a different graph embedding structure from the ST-Transformer network proposed in this patent. Summary of the Invention

[0009] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method and system for supervising abnormal behavior in an interrogation scenario based on the ST-Transformer network.

[0010] According to a method for supervising abnormal behavior in an interrogation scenario based on the ST-Transformer network provided by the present invention, it includes:

[0011] Step S1: Establish a training database and a test database for video samples of the interrogation scenario;

[0012] Step S2: Construct an ST-Transformer network model, and use the training database as input to train the model to obtain a trained ST-Transformer network model;

[0013] Step S3: Obtain and decode the video stream to be detected to get the image of the current frame, and judge whether there is a human body in the image. If not, continue to obtain the next frame of image; if so, extract the human skeleton position in the image and save the human skeleton sequence of the preset frames.

[0014] Step S4: Send the human skeleton sequence to the trained ST-Transformer network model for abnormal behavior detection and analysis;

[0015] Step S5: Save the results of the detection and analysis and the corresponding data to the test database for users to view and supervise.

[0016] Preferably, step S1 includes:

[0017] Step S1.1: Split the existing samples into video segments according to a preset length;

[0018] Step S1.2: Extract the human skeletons in the video segments through a human skeleton extraction algorithm to obtain the human skeleton sequence corresponding to each video segment;

[0019] Step S1.3: Label the normal behaviors and abnormal behaviors included in the human skeleton sequence to obtain an interrogation scene behavior training database and a test database.

[0020] Preferably, establish a human skeleton database, store the human skeletons in the human skeleton database, and use the human skeleton database to label the corresponding categories in the sample management tool;

[0021] The categories include using equipment and passing items;

[0022] The database only includes human skeleton information data.

[0023] Preferably, the structure of the ST-Transformer network model includes four levels, and each level contains a graph embedding module and an ST-Transformer module;

[0024] The graph embedding module performs local feature aggregation on the human skeleton data, including a spatial domain convolutional block, a deformation time domain convolutional block, and an attention layer;

[0025] The ST-Transformer module includes a normalization layer with a residual connection and a spatial attention layer, a normalization layer with a residual connection and a time attention layer, and a normalization layer with a residual connection and a fully connected layer.

[0026] Preferably, step S4 further includes a single-person interrogation and crowd density determination step: count the human skeleton nodes in the video stream to be detected, output the number of people included in the current frame, and perform intelligent determination of single-person interrogation and crowd density based on the number of people.

[0027] According to an abnormal behavior supervision system for an interrogation scene based on the ST-Transformer network provided by the present invention, it includes:

[0028] Module M1: Establish a training database and a test database for the video samples of the interrogation scenario;

[0029] Module M2: Construct an ST-Transformer network model, and use the training database as input to train the model to obtain a trained ST-Transformer network model;

[0030] Module M3: Obtain and decode the video stream to be detected to get the image of the current frame, determine whether there is a human body in the image. If not, continue to obtain the next frame of the image. If so, extract the human body skeleton position in the image and save the human body skeleton sequence of the preset frames;

[0031] Module M4: Send the human body skeleton sequence to the trained ST-Transformer network model for abnormal behavior detection and analysis;

[0032] Module M5: Save the results of the detection and analysis and the corresponding data to the test database for users to view and supervise.

[0033] Preferably, the module M1 includes:

[0034] Module M1.1: Split the existing samples into video segments according to a preset length;

[0035] Module M1.2: Extract the human body skeleton in the video segment through a human body skeleton extraction algorithm to obtain the human body skeleton sequence corresponding to each video segment;

[0036] Module M1.3: Label the normal behaviors and abnormal behaviors included in the human body skeleton sequence to obtain an interrogation scenario behavior training database and a test database.

[0037] Preferably, establish a human body skeleton database, store the human body skeletons in the human body skeleton database, and use the human body skeleton database to label the corresponding categories in the sample management tool;

[0038] The categories include using equipment and passing items;

[0039] The database only includes human body skeleton information data.

[0040] Preferably, the structure of the ST-Transformer network model includes four levels, and each level contains a graph embedding module and an ST-Transformer module;

[0041] The graph embedding module performs local feature aggregation on the human body skeleton data, including a spatial domain convolutional block, a deformation time domain convolutional block, and an attention layer;

[0042] The ST-Transformer module includes a normalization layer and a spatial attention layer with residual connection, a normalization layer and a temporal attention layer with residual connection, and a normalization layer and a fully connected layer with residual connection.

[0043] Preferably, module M4 further includes a single-person interrogation and crowd density determination module: counting the human skeleton nodes in the video stream to be detected, outputting the number of people included in the current frame, and performing single-person interrogation and crowd density intelligent determination according to the number of people.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] 1. The present invention extracts the human skeleton from the original video to generate a video containing only the human skeleton, which can effectively avoid noise interference such as the background environment. At the same time, the original video information in the surveillance video does not need to be stored in the system, protecting the privacy information of the personnel in the interrogation scenario and ensuring the high privacy of the system.

[0046] 2. The present invention also proposes a deformed human skeleton graph embedding module and an ST-Transformer attention network to learn the correlation of joint points, which can more effectively identify the key information in human actions and the algorithm results are better.

[0047] 3. The present invention comprehensively uses the methods of human skeleton and neural network learning to detect abnormal behaviors, constructs a complete set of detection systems, has high recognition accuracy, strong robustness, fast running speed, can output and update the detection results in real time, and is applicable to various scenarios.

[0048] 4. In the ST-Transformer of the present invention, deformable convolutions with bias parameters added are used, which can flexibly adjust the temporal field of view, learn the correlation information in the time domain through additional parameters, and at the same time, each layer of the ST-Transformer module adopts a spatio-temporal separated attention structure, which can better focus on the correlation between the time domain and the spatial domain. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0050] Figure 1 It is a schematic diagram of the working process of the present invention.

[0051] Figure 2 It is a schematic diagram of the process established for the human skeleton extraction and intelligent recognition module in the present invention.

[0052] Figure 3 It is a schematic diagram of the specific structure parameter design of the ST-Transformer network in the present invention.

[0053] Figure 4 Schematic diagram of the human skeleton map embedding module in the present invention

[0054] Figure 5 Schematic diagram of the ST-Transformer module in the present invention Specific implementation manners

[0055] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0056] The present invention focuses on the human skeleton information in video data, and adopts a spatio-temporal attention method to capture the spatial positions and connection relationships between joint points, so as to judge whether there is an abnormal behavior detection method, improving the accuracy and robustness of detection. The human skeleton is extracted from the original video through a pose estimation algorithm, which can ensure the influence of subsequent behavior recognition on the background, environmental noise, etc., and at the same time does not store the original video image, thus ensuring the high privacy of the system. Then, it is trained and recognized through the ST-Transformer network structure.

[0057] Embodiment 1

[0058] According to a method for supervising abnormal behaviors in an interrogation scene based on the ST-Transformer network provided by the present invention, as Figure 1 shown, it includes:

[0059] Step S1: Establish a training database and a test database for video samples in the interrogation scene. Step S1 includes: Step S1.1: Segment video clips from existing samples according to a preset length; Step S1.2: Extract the human skeleton in the video clip through a human skeleton extraction algorithm to obtain a human skeleton sequence corresponding to each video clip; Step S1.3: Label the normal behaviors and abnormal behaviors included in the human skeleton sequence, so as to obtain a behavior training database and a test database for the interrogation scene.

[0060] Specifically, the interrogation scene video sample database contains 1000 samples, which are randomly divided into a training database and a test database according to a ratio of 8:2. It contains a total of 200 normal samples, 400 samples using communication devices, and 400 samples of close contact. Among them, the training database is used to train the model, and the test database is used to test the accuracy of the trained model. Further, a human skeleton database is established, and the human skeletons are stored in the human skeleton database, and the corresponding categories are marked in the sample management tool using the human skeleton database, such as using devices, passing items, etc. The database only includes human skeleton information data.

[0061] Step S2: Construct an ST-Transformer network model, and use the training database as input to train the model to obtain a trained ST-Transformer network model. Specifically, as Figure 3 shown, the structure of the ST-Transformer network model includes four levels, and each level contains a graph embedding module and an ST-Transformer module. Among them, the graph embedding module performs local feature aggregation on the human skeleton data. As Figure 4 shown, it includes a spatial domain convolutional block, a deformable time domain convolutional block, and an attention layer. The deformable time domain convolution can flexibly adjust the time domain view by adding learnable bias parameters, so as to learn the correlation information in the time domain. As Figure 5 shown, the ST-Transformer module includes a normalization layer with residual connection and a spatial attention layer, a normalization layer with residual connection and a time attention layer, and a normalization layer with residual connection and a fully connected layer. At the same time, each layer of the ST-Transformer module adopts a spatio-temporal separation attention structure, which can better focus on the correlation between the time domain and the spatial domain.

[0062] Step S3: Obtain and decode the video stream to be detected to obtain the image of the current frame, and determine whether there is a human body in the image. If not, continue to obtain the next frame of the image; if so, extract the human skeleton position in the image and save the human skeleton sequence of the preset frame.

[0063] Step S4: As Figure 2 shown, send the human skeleton sequence to the trained ST-Transformer network model for abnormal behavior detection and analysis. At the same time, it also includes a single-person interrogation and crowd density determination step: count the human skeleton nodes in the video stream to be detected, output the number of people contained in the current frame, and perform single-person interrogation and crowd density intelligent determination according to the number of people.

[0064] Step S5: Save the results of the detection and analysis and the corresponding data to the test database for users to view and supervise.

[0065] Furthermore, the superiority of the algorithm of the present invention is verified by conducting experiments using the ST-GCN open-source algorithm, and the specific description is as follows:

[0066] Model training and testing are carried out through a human skeleton database under the same interrogation scenario. Among them, the sample size of the training database is 800, and the sample size of the testing database is 200 segments of human skeleton sequences. Through 100 iterations of training, the model is obtained for testing. The final result shows that the accuracy of ST-GCN is 89%, while the accuracy of the ST-Transformer model proposed by the present invention reaches 97% in the test library. Its calculation formula is as follows:

[0067]

[0068] Among them, Accuracy represents the detection result accuracy after model training, N correct represents the sample size of correct classification, N total represents all sample sizes. At the same time, the FLOPs calculation amount of the ST-Transformer model is 1.41G, which is better than 16.3G of the ST-GCN model, indicating the high efficiency of ST-Transformer.

[0069] In summary, through the analysis of the detection results, the present invention has a higher detection rate for the use of communication devices and close contacts. Therefore, in the interrogation scenario, ST-Transformer has better adaptability.

[0070] Embodiment 2

[0071] The present invention also provides an abnormal behavior supervision system for interrogation scenarios based on the ST-Transformer network. Those skilled in the art can implement the abnormal behavior supervision system for interrogation scenarios based on the ST-Transformer network by executing the step process of the abnormal behavior supervision method for interrogation scenarios based on the ST-Transformer network. That is, the abnormal behavior supervision method for interrogation scenarios based on the ST-Transformer network can be understood as the preferred implementation manner of the abnormal behavior supervision system for interrogation scenarios based on the ST-Transformer network.

[0072] An abnormal behavior supervision system for interrogation scenarios based on the ST-Transformer network provided by the present invention includes:

[0073] Module M1: Establish a training database and a test database for video samples of interrogation scenarios. The module M1 includes: Module M1.1: Split video segments of existing samples according to a preset length; Module M1.2: Extract the human skeletons in the video segments through a human skeleton extraction algorithm to obtain a human skeleton sequence corresponding to each video segment; Module M1.3: Label the normal behaviors and abnormal behaviors included in the human skeleton sequence, so as to obtain a training database and a test database for interrogation scenario behaviors. Establish a human skeleton database, store the human skeletons in the human skeleton database, and label the corresponding categories in the sample management tool using the human skeleton database; the categories include using equipment and passing items; the database only includes human skeleton information data.

[0074] Module M2: Construct an ST-Transformer network model, and use the training database as input to train the model to obtain a trained ST-Transformer network model. The structure of the ST-Transformer network model includes four levels, and each level contains a graph embedding module and an ST-Transformer module; the graph embedding module performs local feature aggregation on human skeleton data, including a spatial domain convolutional block, a deformation time domain convolutional block, and an attention layer; the ST-Transformer module includes a normalization layer with residual connection and a spatial attention layer, a normalization layer with residual connection and a time attention layer, and a normalization layer with residual connection and a fully connected layer.

[0075] Module M3: Obtain and decode the video stream to be detected, obtain the image of the current frame, and determine whether there is a human body in the image. If not, continue to obtain the next frame of image; if so, extract the human skeleton position in the image and save the human skeleton sequence of a preset number of frames.

[0076] Module M4: Send the human skeleton sequence to the trained ST-Transformer network model for abnormal behavior detection and analysis. It also includes a single-person interrogation and crowd density determination module: count the human skeleton nodes in the video stream to be detected, output the number of people included in the current frame, and perform intelligent determination of single-person interrogation and crowd density according to the number of people.

[0077] Module M5: Save the results of the detection and analysis and the corresponding data to the test database for users to view and supervise.

[0078] Those skilled in the art know that, in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program code, it is entirely possible to make the systems, devices, and their respective modules provided by the present invention be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. by logically programming the method steps. Therefore, the systems, devices, and their respective modules provided by the present invention can be considered as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the method or the structures within the hardware component.

[0079] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific implementation manners. Those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A method for supervising abnormal behaviors in an interrogation scenario based on the ST-Transformer network, characterized in that, Including: Step S1: Establish a training database and a test database for interrogation scene video samples; Step S2: Construct an ST-Transformer network model, and use the training database as input to train the model to obtain a trained ST-Transformer network model; Step S3: Obtain and decode the video stream to be detected to get the image of the current frame, determine whether there is a human body in the image, if not, continue to obtain the next frame of the image; if so, extract the human skeleton position in the image and save the human skeleton sequence of the preset frames; Step S4: Send the human skeleton sequence to the trained ST-Transformer network model for abnormal behavior detection and analysis; Step S5: Save the results of the detection and analysis and the corresponding data to the test database for users to view and supervise.

2. The abnormal behavior supervision method for interrogation scenarios based on the ST-Transformer network according to claim 1, characterized in that The step S1 includes: Step S1.1: Segment the existing samples into video clips according to a preset length; Step S1.2: Extract the human skeletons in the video clips through a human skeleton extraction algorithm to obtain the human skeleton sequence corresponding to each video clip; Step S1.3: Label the normal behaviors and abnormal behaviors included in the human skeleton sequence to obtain an interrogation scene behavior training database and a test database that both include normal behaviors and abnormal behaviors.

3. The abnormal behavior supervision method for the interrogation scenario based on the ST-Transformer network according to claim 2, characterized in that, Establish a human skeleton database, store the human skeletons in the human skeleton database, and use the human skeleton database to label the corresponding categories in the sample management tool; The categories include using equipment and passing items; The database only includes human skeleton information data.

4. The abnormal behavior supervision method for interrogation scenarios based on the ST-Transformer network according to claim 1, wherein, The structure of the ST-Transformer network model includes four levels, and each level contains a graph embedding module and an ST-Transformer module; The graph embedding module aggregates local features of human skeleton data, including a spatial domain convolutional block, a deformation time domain convolutional block, and an attention layer; The ST-Transformer module includes a normalization layer with residual connection and a spatial attention layer, a normalization layer with residual connection and a temporal attention layer, and a normalization layer with residual connection and a fully connected layer.

5. The abnormal behavior supervision method for interrogation scenarios based on the ST-Transformer network according to claim 1, characterized in that, Step S4 also includes a single-person interrogation and crowd density determination step: count the human skeleton nodes in the video stream to be detected, output the number of people included in the current frame, and perform intelligent determination of single-person interrogation and crowd density according to the number of people.

6. An abnormal behavior supervision system for interrogation scenarios based on the ST-Transformer network, characterized in that, Including: Module M1: Establish a training database and a test database for interrogation scene video samples; Module M2: Construct an ST-Transformer network model, and use the training database as input to train the model to obtain a trained ST-Transformer network model; Module M3: Obtain and decode the video stream to be detected to get the image of the current frame, determine whether there is a human body in the image, if not, continue to obtain the next frame of the image; if so, extract the human skeleton position in the image and save the human skeleton sequence of the preset frames; Module M4: Send the human skeleton sequence to the trained ST-Transformer network model for abnormal behavior detection and analysis; Module M5: Save the results of the detection and analysis and the corresponding data to the test database for users to view and supervise.

7. The abnormal behavior supervision system for the interrogation scenario based on the ST-Transformer network according to claim 6, characterized in that, The said Module M1 includes: Module M1.1: Segment video clips of existing samples according to a preset length; Module M1.2: Extract the human skeletons in the video clips through a human skeleton extraction algorithm to obtain the human skeleton sequence corresponding to each video clip; Module M1.3: Label the normal and abnormal behaviors included in the human skeleton sequence, so as to obtain the interrogation scene behavior training database and test database.

8. The abnormal behavior supervision system for the interrogation scenario based on the ST-Transformer network according to claim 7, characterized in that, Establish a human skeleton database, store the human skeletons in the human skeleton database, and label the corresponding categories in the sample management tool using the human skeleton database; The said categories include using equipment and passing items; The database only includes human skeleton information data.

9. The abnormal behavior supervision system for the interrogation scenario based on the ST-Transformer network according to claim 6, characterized in that, The structure of the said ST-Transformer network model includes four levels, and each level contains a graph embedding module and an ST-Transformer module; The said graph embedding module performs local feature aggregation on human skeleton data, including a spatial domain convolutional block, a deformation time domain convolutional block, and an attention layer; The said ST-Transformer module includes a normalization layer with residual connection and a spatial attention layer, a normalization layer with residual connection and a time attention layer, and a normalization layer with residual connection and a fully connected layer.

10. The abnormal behavior supervision system for the interrogation scenario based on the ST-Transformer network according to claim 6, characterized in that, Module M4 also includes a single-person interrogation and crowd density determination module: Count the human skeleton nodes in the video stream to be detected, output the number of people included in the current frame, and perform intelligent determination of single-person interrogation and crowd density according to the number of people.

Citation Information

Patent Citations

  • Abnormal behavior recognition method based on full modality

    CN113076772A

  • Abnormal behavior detection method and system based on spatial diagram convolutional network

    CN113269103A

  • Abnormal behavior intelligent detection method and system based on similarity graph neural network

    CN113065515A

  • Integration platform for multi-network integration of service platforms

    US20180359132A1