Campus violent behavior detection method and system
By adopting a comprehensive detection system that integrates multiple features in campus monitoring, combined with dynamic analysis of motion trajectory and overall understanding of video content, the problems of insufficient feature utilization, difficulty in taking into account real-time and accuracy, and poor adaptability in the existing technology are solved, and efficient detection and early warning of violent behaviors in campuses are achieved.
Patent Information
- Application Number
- CN202411970631.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has problems such as missed and missed detection caused by insufficient utilization of features, difficulty in taking into account real-time and accuracy, and poor adaptability in violent behavior detection in campus monitoring.
A comprehensive detection system with multiple features fusion is adopted, combining dynamic analysis of motion trajectory and overall understanding of video content, target detection is used using the YOLOv8 model, BoT-SORT algorithm for target tracking, long-term and short-term memory model analyzes motion trajectory, and video understanding model TSM captures timing information, and performs feature fusion for detection.
It realizes early warning and timely response to violent behaviors in campus, takes into account detection accuracy, calculation complexity and real-timeness, and improves adaptability in campus scenarios.
Smart Images

Figure CN120014536A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and in particular to a method and system for detecting campus violence. Background Art
[0002] In recent years, school bullying incidents have occurred frequently, attracting widespread attention from all walks of life. School bullying has caused great harm to the victims, but they are often unwilling to take the initiative to report to their parents and teachers due to self-esteem and fear. Video surveillance is an important means of security prevention, but manual verification is inefficient.
[0003] In recent years, with the rapid development of artificial intelligence technology, it has become a new trend to combine deep learning algorithms with surveillance videos to automatically complete detection and recognition tasks. For violent behavior detection, existing technologies mainly focus on methods based on target detection, skeleton point analysis, and video understanding.
[0004] The target detection-based violent behavior detection method regards fighting behavior as a target category and uses mature target detection algorithms (such as the YOLO series or Faster R-CNN) to classify and detect each frame of the video to identify whether there is fighting behavior. For example, the existing invention patent application document "A method, device, equipment and medium for violence detection based on neural network" with publication number CN112287754A, the existing method includes: sending the video to be detected to a violence detection model pre-deployed on the cloud platform to determine the security level; judging whether the video to be detected is a violent event according to the security level; if it is determined that the video to be detected contains a violent event, issuing a corresponding warning message according to the security level. The violence detection model is constructed based on a 3D CNN network model and a 3D RNN network model, wherein the CNN network model is used to extract the spatial features of the image; and the RNN network model is used to extract the temporal features of the image. Although this method is relatively mature and has strong real-time performance, due to the lack of utilization of temporal information, the detection method based on a single frame has the risk of missed detection when identifying continuous violent behavior, and has poor adaptability to complex scenes such as occlusion and illumination changes.
[0005] The detection method based on skeleton points extracts the key point information of human skeleton through frameworks such as OpenPose or MediaPipe, and judges violent behavior in combination with logical rules. For example, fighting behavior is identified by analyzing the amplitude, direction and other characteristics of limb movements. In the existing invention patent application document "Violence Detection System and Detection Method Based on Distributed Monitoring" with publication number CN113269076A, M surveillance cameras process M surveillance camera captured images respectively through the improved COCO model embedded in M embedded devices, and the specific method for obtaining the processed surveillance camera captured image data is: using M embedded devices to extract the key points of human skeleton from the M surveillance camera captured images obtained in step 1, and then using the improved OpenPose model to extract human skeleton points to form a human skeleton model, forming a video stream containing only the skeleton model; the video stream containing only the skeleton model is used as the processed surveillance camera captured image data. This type of existing method has certain advantages in reducing background interference and can focus more on the behavioral characteristics of the characters. However, this method relies heavily on video clarity and skeleton point extraction accuracy, and its performance will be significantly reduced in low-quality videos or occluded scenes. In addition, rule-based judgment methods lack flexibility and adaptability, and are difficult to deal with diverse and complex violent behaviors.
[0006] The three-dimensional detection method based on spatiotemporal feature extraction uses the temporal characteristics of the video to capture the dynamic characteristics of violent behavior through a deep learning model (such as a three-dimensional convolutional network C3D or a temporal convolutional network TCN) to analyze the video content as a whole. This method can make full use of the temporal information of the behavior and has a strong ability to capture complex actions. It is especially suitable for analyzing continuous or short-term violent behavior. However, this method has a high computational cost and is difficult to achieve real-time processing on resource-constrained devices. In addition, it has a strong dependence on large-scale annotated datasets, and when applied to specific scenarios (such as campus monitoring), the generalization ability of the model may be limited.
[0007] The application of existing violent behavior detection technology in actual campus monitoring still has certain shortcomings. On the one hand, the use of a single feature limits the comprehensive characterization of violent behavior in complex dynamic scenes; on the other hand, it is difficult to balance the real-time and accuracy of detection, especially when resources are limited. In addition, most of the currently available data sets are public scenes, lacking specific data for campus environments, resulting in poor adaptability of the model in campus monitoring scenarios.
[0008] In summary, the existing technology has technical problems such as missed detection and false detection due to insufficient feature utilization, difficulty in balancing real-time and accuracy, and poor adaptability in campus monitoring scenarios. Summary of the invention
[0009] The technical problem to be solved by the present invention is: how to solve the technical problems in the prior art such as missed detection and false detection caused by insufficient feature utilization, difficulty in balancing real-time performance and accuracy, and poor adaptability in campus monitoring scenarios.
[0010] The present invention adopts the following technical solutions to solve the above technical problems: A method for detecting campus violence includes:
[0011] S1. Build and combine public data sets and self-built campus scene data sets, filter out campus environment video clips, and build campus scene violence behavior data sets based on campus environment video clips;
[0012] S2. Detect violent behavior in campus surveillance videos based on motion trajectory analysis; construct and use the YOLOv8 model as a detector to detect pedestrian targets in surveillance videos from the campus scene violent behavior dataset; combine the tracking algorithm to track pedestrian targets in real time and extract motion trajectories; perform behavioral analysis on pedestrian targets, analyze the tracked motion trajectories through the long-short-term memory model, extract abnormal behavior pattern features, and use them to determine the behavior of pedestrian targets for early warning and response to violent behavior;
[0013] S3. Set up and use the video understanding model TSM to perform video understanding operations, capture the temporal information of the video sequence in the campus scene violence data set, perform time offset operations on the channel features in the preset time dimension, and exchange local time information in the input feature map of each time step to extract strong temporal dependency features and perceive the dynamic behavior of pedestrian targets; perform feature fusion operations on the output data of the video understanding model TSM and the YOLOv8 model to obtain a fused feature map, which is used to detect violence in campus surveillance line videos.
[0014] The present invention provides two innovative technical solutions for campus safety management based on the dynamic analysis of motion trajectories and the overall understanding of video content by constructing a comprehensive detection system that can integrate multiple features, meeting the detection of violent behavior in campus surveillance videos, thereby achieving early warning and timely response to violent behavior on campus. The present invention designs a method that integrates target detection and time series features, taking into account detection accuracy, computational complexity, and real-time performance, while making full use of customized data sets for campus scenes, which can meet the actual needs of campus safety management for violent behavior detection.
[0015] The present invention adopts TSM, which is an efficient video understanding model based on time dimension feature offset, which can effectively capture the timing information in the video sequence, thereby enhancing the perception of dynamic behavior. TSM offsets the features of some channels in the time dimension without increasing the computational complexity through the time offset operation, thereby achieving effective information propagation in the time domain. This design enables the model to fully explore the temporal dependencies in the video sequence and makes the flow of temporal information more efficient, thereby achieving good temporal feature extraction effects without requiring a lot of calculations.
[0016] In a more specific technical solution, S1 includes:
[0017] S11, screening out violent videos, non-violent videos, and campus scene videos to construct a public data set;
[0018] S12. Use pre-set collection equipment to collect surveillance videos of violent behaviors to build a self-built campus scene dataset;
[0019] S13. Refine the campus violence behavior dataset.
[0020] In order to meet the actual needs of campus violence detection algorithms, the present invention combines multiple public data sets with self-built campus scene data sets, and screens and supplements video clips suitable for campus environments in a targeted manner. Finally, a high-quality campus violence data set is designed and constructed, and the data is processed in a refined manner to ensure data diversity and high relevance to practical applications.
[0021] In a more specific technical solution, S13 includes:
[0022] S131, dividing the videos in the campus scene violence data set to convert them into frame sequences for time series analysis;
[0023] S132, using a preset annotation tool to annotate the frame sequence frame by frame to obtain an annotated frame sequence;
[0024] S133: Perform enhancement processing on the marked frame sequence.
[0025] In a more specific technical solution, in S2, the BoT-SORT algorithm is used to perform IoU matching operations and re-identification ReID feature fusion operations to track the motion trajectory of each pedestrian target in a multi-target situation;
[0026] In the BoT-SORT algorithm, all pedestrian targets in the video frame are detected through the YOLOv8 model;
[0027] According to the intersection over union (IoU) matching mechanism of pedestrian targets, the targets are tracked between consecutive frames, and the ReID feature vector is obtained and combined to perform an ID assignment operation for each pedestrian target, thereby continuously tracking the motion trajectory of the pedestrian target in the video stream.
[0028] In the process of violent behavior identification, motion trajectory data fusion is performed to analyze and obtain behavioral feature difference data and perform feature extraction operations.
[0029] The feature extraction operation includes: adjusting the aspect ratio of the detection rectangle; integrating the trajectory information of all IDs in the original video and the aspect ratio of the detection rectangle of the ID into comprehensive data.
[0030] In a more specific technical solution, in S2, the YOLOv8 model includes: a channel attention module and a spatial attention module; which are used to focus on processing the key features of video frames in the campus scene violence data set.
[0031] The YOLOv8 used in the present invention as the target detection model has achieved a good balance in detection accuracy, resource utilization and real-time performance, and is suitable for detecting violent behavior in campus monitoring environments. The present invention introduces channel attention and spatial attention mechanisms into the model, allowing the model to focus more on key features in the video frame.
[0032] Using the channel attention module, the channel weight vector is generated through global average pooling to adjust the attention of the YOLOv8 model to different feature channels;
[0033] Using the spatial attention module, we generate a spatial weight map through maximum pooling combined with convolution operations to optimize the YOLOv8 model's attention to different locations in the video frame and focus on areas containing violent behavior.
[0034] Use the frame-level annotation dataset to train the YOLOv8 model.
[0035] During the training process, the model uses a frame-level annotated data set to ensure that in complex environments with multiple targets, overlapping people, and scene occlusion, the model can still effectively locate and identify areas and actions with violent characteristics. YOLOv8 combined with the attention mechanism not only performs well in complex scenes, but also has good real-time processing capabilities, enabling it to quickly respond to and detect potential violent behaviors in a campus environment, ensuring the accuracy, robustness, and practicality of the system.
[0036] In a more specific technical solution, in S2, the BoT-SORT algorithm is used to track the target, perform IoU matching operations, and re-identify ReID feature fusion operations to track the motion trajectory of each pedestrian target in the case of multiple targets;
[0037] In the BoT-SORT algorithm, all pedestrian targets in the video frame are detected through the YOLOv8 model;
[0038] According to the IoU matching mechanism of pedestrian targets, the targets are tracked between consecutive frames, and the ReID feature vector is obtained and combined to assign an ID to each pedestrian target, so as to continuously track the motion trajectory of pedestrian targets in the video stream.
[0039] In the process of violent behavior identification, motion trajectory data fusion is performed to analyze the behavior feature difference data and perform feature extraction operations, where the feature extraction operations include: adjusting the aspect ratio of the detection rectangle; integrating the trajectory information of all IDs in the original video and the aspect ratio of the detection rectangle of the ID into comprehensive data.
[0040] The present invention adopts the BoT-SORT algorithm, which combines IoU matching with ReID (re-identification) feature fusion to ensure that the motion trajectory of each target can be continuously and accurately tracked in the case of multiple targets. The campus surveillance video violence detection algorithm proposed by the present invention, which is based on the integration of target detection, target tracking and time series analysis, significantly improves the detection accuracy by deeply analyzing the motion trajectory of pedestrians in the video.
[0041] In a more specific technical solution, in S2, a long short-term memory network model LSTM is constructed, wherein the long short-term memory network model LSTM includes: an LSTM layer, a Bi-LSTM layer, an adaptive average pooling layer and a fully connected layer.
[0042] In a more specific technical solution, input tensor parameters are set; LSTM layers are set and used to obtain and capture time dependencies in sequence data based on hidden states and cell states, and obtain LSTM output data;
[0043] Set up and use the Bi-LSTM layer to process the forward time series and the backward time series at the same time to obtain the Bi-LSTM output data;
[0044] Concatenate the LSTM output data and the Bi-LSTM output data to obtain the forward and bidirectional time series information feature vectors, and extract the hidden state of the last time step based on them;
[0045] Set and use an adaptive average pooling layer to perform a pooling operation on the target motion features in the hidden state of the last time step to obtain a pooled feature vector;
[0046] Set up and use the fully connected layer to process the pooled feature vector, integrate the time features extracted by the LSTM layer and the Bi-LSTM layer, and obtain the output data of the fully connected layer;
[0047] Using the sigmoid activation function, the output data of the fully connected layer is mapped to the preset interval, and the probability score of violent behavior is predicted.
[0048] In a more specific technical solution, the feature fusion operation includes:
[0049] S31, concatenate the 2D features output by the FPN layer P2 of the YOLOv8 model and the 3D features output by the backbone layer layer2 of the time-frequency understanding model TSM in the preset channel dimension to obtain tensor concatenated features;
[0050] S32, pass the tensor concatenation features through two convolutional layers to obtain the feature tensor F1;
[0051] S33, reshape the feature tensor F1 into a first three-dimensional feature FQ for matrix operation;
[0052] S34, performing a transposition operation on the first three-dimensional feature FQ to obtain a transposed feature FK;
[0053] S35, performing softmax normalization on the dot product of the first three-dimensional feature FQ and the transposed feature FK to generate an attention map Fat;
[0054] S36, reshape the feature tensor F1 to obtain the first three-dimensional feature FV, and multiply it with the attention map Fat to obtain a fused feature map;
[0055] S37, add the fused feature map to the feature tensor F1; after two convolution operations, complete the feature dimension recovery operation and absorb the effective information of the 2D feature.
[0056] In a more specific technical solution, a school violence detection system includes:
[0057] The violent behavior dataset construction module is used to construct and combine public datasets and self-built campus scene datasets, filter out campus environment video clips, and construct campus scene violent behavior datasets based on campus environment video clips;
[0058] The motion trajectory analysis module is used to detect violent behavior in campus surveillance videos based on motion trajectory analysis. The YOLOv8 model is constructed and used as a detector to detect pedestrian targets in surveillance videos from the campus scene violent behavior dataset. The pedestrian targets are tracked in real time with the tracking algorithm to extract the motion trajectory. The pedestrian targets are analyzed for behavior, and the motion trajectory obtained by tracking is analyzed by the long-short-term memory model to extract abnormal behavior pattern features, which are used to determine the behavior of the pedestrian targets for violent behavior warning and response. The motion trajectory analysis module is connected to the violent behavior dataset construction module.
[0059] The video understanding module is used to set up and use the video understanding model TSM to perform video understanding operations, capture the temporal information of the video sequence in the campus scene violence data set, perform time offset operations on the channel features in the preset time dimension, and exchange local time information in the input feature map of each time step to extract strong temporal dependency features and perceive the dynamic behavior of pedestrian targets; perform feature fusion operations on the output data of the video understanding model TSM and the YOLOv8 model to obtain a fused feature map, which is used to detect violence in campus surveillance line videos. The video understanding module is connected to the violent behavior data set construction module.
[0060] Compared with the prior art, the present invention has the following advantages:
[0061] The present invention provides two innovative technical solutions for campus safety management based on the dynamic analysis of motion trajectories and the overall understanding of video content by constructing a comprehensive detection system that can integrate multiple features, meeting the detection of violent behavior in campus surveillance videos, thereby achieving early warning and timely response to violent behavior on campus. The present invention designs a method that integrates target detection and time series features, taking into account detection accuracy, computational complexity, and real-time performance, while making full use of customized data sets for campus scenes, which can meet the actual needs of campus safety management for violent behavior detection.
[0062] The present invention adopts TSM, which is an efficient video understanding model based on time dimension feature offset, which can effectively capture the timing information in the video sequence, thereby enhancing the perception of dynamic behavior. TSM offsets the features of some channels in the time dimension without increasing the computational complexity through the time offset operation, thereby achieving effective information propagation in the time domain. This design enables the model to fully explore the temporal dependencies in the video sequence and makes the flow of temporal information more efficient, thereby achieving good temporal feature extraction effects without requiring a lot of calculations.
[0063] In order to meet the actual needs of campus violence detection algorithms, the present invention combines multiple public data sets with self-built campus scene data sets, and screens and supplements video clips suitable for campus environments in a targeted manner. Finally, a high-quality campus violence data set is designed and constructed, and the data is processed in a refined manner to ensure data diversity and high relevance to practical applications.
[0064] The YOLOv8 used in the present invention as the target detection model has achieved a good balance in detection accuracy, resource utilization and real-time performance, and is suitable for detecting violent behavior in campus monitoring environments. The present invention introduces channel attention and spatial attention mechanisms into the model, allowing the model to focus more on key features in the video frame.
[0065] The present invention adopts the BoT-SORT algorithm, which combines IoU matching with ReID (re-identification) feature fusion to ensure that the motion trajectory of each target can be continuously and accurately tracked in the case of multiple targets. The campus surveillance video violence detection algorithm proposed by the present invention, which is based on the integration of target detection, target tracking and time series analysis, significantly improves the detection accuracy by deeply analyzing the motion trajectory of pedestrians in the video.
[0066] The present invention solves the technical problems existing in the prior art, such as missed detection and false detection caused by insufficient feature utilization, difficulty in balancing real-time performance and accuracy, and poor adaptability in campus monitoring scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 A schematic diagram of the basic steps of a method for detecting campus violence in Example 1 of the present invention;
[0068] Figure 2 This is a schematic diagram of a campus violence behavior dataset according to Example 1 of the present invention;
[0069] Figure 3 This is a schematic diagram of data stream processing of a campus surveillance video violence behavior detection model based on motion trajectory analysis according to Example 1 of the present invention;
[0070] Figure 4 This is a schematic diagram of data stream processing for detecting violent behavior in campus surveillance videos based on video content understanding according to Embodiment 1 of the present invention;
[0071] Figure 5 This is a schematic diagram of feature fusion data stream processing in Example 1 of the present invention;
[0072] Figure 6 Schematic diagram of attention fusion of layer3, layer4 of TSM and YOLOv8 model in embodiment 2 of the present invention. DETAILED DESCRIPTION
[0073] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0074] Example 1
[0075] like Figure 1 As shown, a campus violence detection method provided by the present invention comprises the following basic steps:
[0076] S1. Construct and process a violent behavior dataset in campus scenarios. The violent behavior dataset includes: public dataset and self-built dataset;
[0077] like Figure 2 As shown, in this embodiment, in order to meet the actual needs of the campus scene violence detection algorithm, the present invention combines multiple public data sets with self-built campus scene data sets, and specifically screens and supplements video clips suitable for the campus environment, and finally designs and constructs a high-quality campus scene violence data set. The data is processed in a refined manner to ensure the diversity of the data and the high relevance of practical applications.
[0078] In the process of constructing the public data set of this embodiment, the advantages of multiple existing data sets are comprehensively utilized to screen out clips related to campus scenes to improve the generalization ability of the model. For example, some violent and non-violent videos in ice hockey games are selected from the Hockey Fight data set to increase the diversity of samples; videos of typical campus scenes such as playgrounds and corridors are selected from the RWF-2000 data set to enhance the adaptability of the model to the actual campus environment; the UCF-Crime data set provides a wide range of abnormal behavior samples, and supplements the coverage of the data set by screening fights and non-violent behavior clips; the real surveillance perspective video in the Surveillance Camera Fight data set provides the training model with scene support that is closer to actual application. The fusion of these public data serves as a basis to ensure the richness and representativeness of the data set.
[0079] In the process of constructing the self-built dataset in this embodiment, the actual monitoring scene was simulated on campus. To ensure the authenticity and diversity of the samples, the data was captured by combining the perspective of a drone with a fixed camera, and videos of violent behaviors (such as fighting, pulling, kicking) and non-violent behaviors (such as standing, walking, riding a bicycle, and trotting) in environments such as playgrounds, corridors, and classrooms were collected.
[0080] During the data processing process, the long video is first divided into short clips of fixed length using the editing tool and converted into a frame sequence to meet the input requirements of the timing analysis model. The professional labeling tool LabelImg is then used to accurately label the video frame by frame, including the bounding box position and behavior category of the target person, to ensure that the labeling results are accurate and consistent. To enhance the robustness of the model, data augmentation technology is used to process the clips, including rotation, cropping, brightness adjustment, and mirror flipping operations to expand sample diversity and improve the model's adaptability to lighting changes and noise interference. After processing, the data is converted to the standard format required by YOLOv8 to ensure a seamless training process.
[0081] The final dataset contains 1,980 videos of violent behavior and 1,871 videos of non-violent behavior, covering a variety of campus scenarios, fully meeting the model's requirements for data diversity.
[0082] S2. Detect violent behavior in campus surveillance videos based on motion trajectory analysis; construct and use the YOLOv8 model as a detector to detect pedestrian targets in surveillance videos from the campus scene violent behavior dataset; combine the tracking algorithm to track pedestrian targets in real time and extract motion trajectories; perform behavioral analysis on pedestrian targets, analyze the tracked motion trajectories through the long-short-term memory model, extract abnormal behavior pattern features, and use them to determine the behavior of pedestrian targets for early warning and response to violent behavior;
[0083] like Figure 3 As shown in the figure, the traditional behavior detection framework includes: video input, frame-level segmentation, data preprocessing, feature extraction and training, and behavior recognition. It is mainly based on frame-level image processing to obtain behavior features. Based on the guidance of the traditional behavior detection framework and the analysis of advanced technologies, this paper proposes a method for detecting violent behavior in campus surveillance videos based on dynamic analysis of motion trajectories. Figure 3 .
[0084] In this embodiment, the offline video violence detection model for campus monitoring first uses the YOLO model as a detector to detect only pedestrian targets in the surveillance video, and then combines the tracking algorithm to track the characters in the video in real time, extracting the motion trajectory containing spatiotemporal information, thereby enhancing the understanding of individual behavior and providing data support for behavior pattern recognition. In the behavior analysis stage, the motion trajectory obtained by tracking is analyzed by the long short-term memory model to extract the characteristics of abnormal behavior patterns, and the behavior is judged, and finally warning and responding to potential violent behavior.
[0085] In this embodiment, YOLOv8 is used as the target detection model because it achieves a good balance in detection accuracy, resource utilization and real-time performance, and is particularly suitable for violent behavior detection in campus monitoring environments. In order to further improve the model's sensitivity to violent behavior and detection effect, the present invention introduces channel attention and spatial attention mechanisms into the model, so that the model can focus more on key features in the video frame. The channel attention module generates a channel weight vector through global average pooling, so that the model can automatically adjust the attention to different feature channels and effectively enhance the perception of specific action features; the spatial attention module generates a spatial weight map through maximum pooling combined with convolution operations, further optimizing the model's attention to different positions in the video frame, so that it can more accurately focus on key areas containing violent behavior, thereby achieving dual optimization at the spatial and feature levels, so that the model can robustly detect and analyze violent behavior in complex monitoring scenarios. During the training process, the model uses a frame-level annotated data set to ensure that in a complex environment with multiple targets, multiple people overlapping, and scene occlusion, the model can still effectively locate and identify areas and actions with violent characteristics. YOLOv8 combined with the attention mechanism not only performs well in complex scenarios, but also has good real-time processing capabilities, enabling it to quickly respond to and detect potential violent behaviors in campus environments, ensuring the accuracy, robustness and practicality of the system.
[0086] In the target tracking part of this embodiment, the BoT-SORT algorithm is used, which combines IoU matching and ReID (re-identification) feature fusion to ensure that the motion trajectory of each target can be continuously and accurately tracked in the case of multiple targets. Specifically, the BoT-SORT process first detects all targets in the video frame through YOLOv8, and then tracks the target between consecutive frames based on the target's IoU (intersection over union) matching mechanism, and further combines the ReID feature vector (a high-dimensional feature vector generated by a deep neural network to characterize the uniqueness of the target) to ensure that the same target can be correctly identified in the event of target overlap, occlusion or other interference. In this way, the system is able to assign a unique ID to each detected pedestrian and continuously track the motion trajectory of the ID in the video stream.
[0087] In the violent behavior identification process of this embodiment, in order to improve the accuracy and robustness of detection, the present invention proposes a fusion method of motion trajectory data. The core of this method is to deeply analyze the feature differences between fighting behavior and non-fighting behavior, and propose targeted feature extraction measures accordingly.
[0088] Non-fighting behavior characteristics: In non-fighting videos, the pedestrian’s movement trajectory is relatively stable, and the trajectory of the coordinates of the upper left corner of the tracking box fluctuates slightly, indicating that the pedestrian is in a stable movement state.
[0089] Fighting behavior characteristics: In fighting scenes, the motion trajectory shows significant anomalies. For example, when two or three people fight, the X and Y values fluctuate violently within a large range, and turning points appear frequently in the trajectory. In complex situations where multiple people are watching or fighting, although the trajectories of the bystanders are relatively clear, the violent movements of the fighters often cause a large amount of tracking frame information to be lost, resulting in more complex trajectory fluctuations.
[0090] In order to more accurately identify fighting behavior, the single location feature needs to be improved:
[0091] Add aspect ratio recording: In addition to recording the XY coordinates, the aspect ratio of the detection rectangle is also recorded. This change can effectively capture characteristic movements such as the waving of the arms of the fighters, providing an important basis for identification.
[0092] Comprehensive ID trajectory data: The trajectory information of all IDs in the video and the aspect ratio of their detection rectangles are integrated into a comprehensive data. With this method, we can more comprehensively consider the movement of each pedestrian in the video and improve the accuracy of recognition. For example, in a fight video, the trajectory information and the aspect ratio of the rectangles of ID1 and ID2 are combined into a data and input into the long short-term memory (LSTM) behavior recognition model for analysis.
[0093] Long Short-Term Memory (LSTM) is a recurrent neural network (RNN) architecture in the field of deep learning, which is suitable for processing long-term dependency problems in time series data. Unlike standard feedforward neural networks, LSTM has feedback connections, which can effectively handle time dependencies in sequence data and solve the gradient vanishing or exploding problems common in traditional RNNs. LSTM introduces memory cells, each of which includes an input gate, a forget gate, and an output gate. These gates control the flow of information in the memory cell, allowing the model to retain key information in longer time series. Bi-LSTM (bidirectional LSTM) is an improvement on LSTM, which not only captures forward time series information, but also handles reverse time dependencies, and then performs better when processing complex time series data.
[0094] The process of building an LSTM model is as follows:
[0095] Input tensor: The model receives a 4D input tensor of shape (batch_size, num_persons, seq_len, input_size), where batch_size is the number of videos in the batch, num_persons is the maximum number of individuals tracked in the video, seq_len is the length of the frame sequence for each video, and input_size is the number of input features per frame.
[0096] ·LSTM and Bi-LSTM layers: LSTM layer: LSTM networks are specifically designed to capture temporal dependencies in sequence data. Through hidden states and cell states, LSTM is able to remember key information in long time series, and is particularly suitable for processing forward dependencies of motion trajectories. The output of LSTM is the hidden state of each time step, which retains the information of that time step. Bi-LSTM layer: Bidirectional LSTM (Bi-LSTM) processes both forward and backward time series at the same time. This means that it not only considers the past information of the current frame, but also considers future information through reverse processing. In this way, Bi-LSTM captures more comprehensive temporal dependencies, especially in complex motion patterns (such as sudden motion reversals or changes in direction). Combination of LSTM and Bi-LSTM: When the two are combined, the output of LSTM is concatenated with the output of Bi-LSTM to form a feature vector containing forward and bidirectional temporal information. This combination can provide a more detailed understanding of the changes in the target's motion trajectory.
[0097] Pooling layer: After the input passes through the LSTM and Bi-LSTM layers, the hidden state of the last time step in the sequence is extracted (or other strategies such as averaging the hidden states of all time steps). This state represents the motion characteristics of the individual in the entire sequence. For the motion characteristics of each target, an adaptive average pooling layer is used for pooling. The purpose of this is to cope with the changes in the number of individuals that may appear in the video and ensure that the model can handle different numbers of individuals.
[0098] Fully connected layer: The pooled feature vector is input to the fully connected layer. This fully connected layer further processes the features and integrates the temporal features extracted by LSTM and Bi-LSTM. Using the sigmoid activation function, the output is mapped to the [0,1] interval to generate a score for predicting the probability of violent behavior.
[0099] Through the design of this model, we can effectively utilize the information in the time series, deal with the movement trajectories of multiple individuals at the same time, and achieve efficient detection of violent behavior. The ability of the LSTM network to capture time series information and the ability of the adaptive pooling layer to handle an uncertain number of individuals make the model have good generalization and robustness in practical applications.
[0100] S3. Based on the overall understanding of the video content, campus surveillance line video violence detection is performed; the video understanding model TSM is set and used to perform video understanding operations, capture the temporal information of the video sequence in the campus scene violence data set, perform time offset operations on the channel features in the preset time dimension, and perform local time information exchange in the input feature map of each time step to extract strong temporal dependency features and perceive the dynamic behavior of pedestrian targets; feature fusion operations are performed on the output data of the video understanding model TSM and the YOLOv8 model to obtain a fused feature map, which is used to detect campus surveillance line video violence.
[0101] Traditional violent behavior detection methods usually rely on frame-level segmentation and single-frame feature extraction. This local feature-based method is easily disturbed by background noise and is difficult to accurately capture the overall semantic information of the video in complex scenes. In order to make up for this deficiency, the present invention proposes a campus surveillance video violent behavior detection method from the perspective of overall understanding of video content. By integrating spatiotemporal features, the video content is comprehensively analyzed to improve the accuracy and robustness of violent behavior recognition.
[0102] In this embodiment, in order to further improve the recognition of character behavior and extract accurate action features, the present invention uses a video understanding model TSM. TSM is an efficient video understanding model based on time dimension feature offset, which can effectively capture the timing information in the video sequence, thereby enhancing the perception of dynamic behavior. Compared with the traditional 3D convolutional network, the core advantage of TSM is that it uses the "time offset" operation to offset the features of some channels in the time dimension without increasing the computational complexity, thereby achieving effective information propagation in the time domain. This design enables the model to fully explore the temporal dependencies in the video sequence and makes the flow of temporal information more efficient, thereby achieving good temporal feature extraction effects without the need for a lot of calculations. The backbone network of the model further enhances the model's ability to capture action features by loading pre-trained weights on large video datasets (such as Kinetics, YouTube-8M, etc.). In specific applications, TSM uses time offset operations to exchange local temporal information in the input feature map at each time step, thereby extracting features with strong temporal dependencies, reducing the large amount of computing resources required by traditional 3D convolutional networks, and improving the efficiency and effectiveness of the model when processing long-time series video data.
[0103] In order to further improve the accuracy and robustness of violent behavior detection, this paper proposes to fuse the backbone network features of YOLO with the backbone network features of the video understanding model, such as Figure 4As shown in the figure. Feature fusion can be performed in one or more layers of the pyramid network, and the specific number of layers depends on the task requirements and the computing power of the model. The fusion process can essentially be regarded as an application of an attention mechanism, which allows the model to focus more on the characteristics of the characters in the video and reduce background interference, thereby improving the recognition accuracy of fighting behavior. In this process, the YOLO model is responsible for target detection, detecting the human targets in the video and providing accurate bounding boxes, while the video understanding model extracts sequence information from the time dimension to capture the dynamic interaction between the characters. After fusing the target information detected by the YOLO model with the temporal features extracted by the video understanding model, a richer feature representation is formed. These fused features are further used for classification prediction to determine whether there is fighting in the video. Finally, in the inference stage, when the model detects fighting behavior, it will output the corresponding results, timely warning, and ensure rapid response and handling of potential violent behavior.
[0104] The core goal of feature fusion is to organically integrate the 2D spatial features of the YOLOv8 target detection network and the 3D temporal features of video understanding models such as TSM to improve the accuracy and robustness of the model in the task of violent behavior detection. Traditional 2D features can effectively capture spatial information in video frames, such as the position, shape, and posture of characters; while 3D features have good capture capabilities in the temporal dimension and can characterize dynamic behavioral characteristics, including the continuity and temporal changes of actions. Through this fusion of 2D and 3D features, the model can more accurately identify the interactions between characters and their dynamic changes, thereby significantly improving the sensitivity of violent behavior detection, reducing background interference, and strengthening the focus on violent behavior. This fusion helps to improve the recognition accuracy of the system in complex video surveillance scenarios, enabling it to better cope with the real-time detection needs of violent behaviors such as fighting and attacking.
[0105] The feature fusion operation is as follows: Figure 5 As shown in the figure, it mainly includes the following aspects. First, the 2D features (64×192×28×28) from the YOLOv8 backbone network and the 3D features (64×512×28×28) from the TSM backbone network are concatenated in the channel dimension to obtain a 64×704×28×28 tensor. The concatenation operation enables the model to combine 2D and 3D information, thereby enhancing the expressiveness of the features. Next, the concatenated features pass through two convolutional layers to obtain the feature tensor F1. The convolution operation can further extract effective features and appropriately reduce the feature dimension to control the computational overhead. Subsequently, F1 is reshaped into a three-dimensional feature FQ (b, c, h*w) to facilitate matrix operations and improve computational efficiency.
[0106] Then, FQ is transposed to obtain a new feature FK, ensuring that the dimensions of the query (Q) and key (K) in the attention mechanism match. In the attention mechanism, the dot product of FQ and FK is normalized by softmax to generate the attention map Fat, which dynamically focuses on the important parts of the input features, thereby highlighting the model's attention to the character's actions and reducing its dependence on the background area. At the same time, F1 is reshaped into a three-dimensional feature FV again and multiplied with the attention map Fat to obtain a fused feature map. At this time, the fused feature is added to F1 to maintain the consistency of the feature. Finally, after two convolution operations, the dimensions of the original 3D feature (64×512×28×28) are restored to ensure that the effective information from the 2D feature is absorbed while maintaining the accuracy of the 3D feature.
[0107] This fusion method cleverly introduces the attention mechanism, focusing on key action areas, effectively reducing the impact of background interference, and significantly improving the performance of the model in violent behavior detection. During the reasoning process, the model can detect and identify the behavior of people in the video in real time, and output an alarm signal in a timely manner when it is confirmed that violent behavior has occurred. This design not only improves the model's accuracy in identifying violent behavior in complex scenarios, but also enhances the system's real-time warning capabilities, providing accurate and efficient security protection support for key monitoring sites such as campuses.
[0108] Example 2
[0109] In this embodiment, the test was conducted on the same data set Surveillance Camera Fight. Compared with other violent behavior detection algorithms that only combine time series analysis models (such as VGG16+LSTM with a detection accuracy of 62%, VGG16+BiLSTM with 45%, Xception+LSTM with 60%, Xception+BiLSTM with 63.3%, Xception+BiLSTM+attention with 69%, etc.), this algorithm achieved a detection accuracy of 81.63%, reaching the current best level, fully demonstrating its robustness and superiority in complex scenarios. The ablation experiment of the violent behavior detection algorithm of campus surveillance videos based on video content understanding showed that the accuracy of the control group using the TSM model alone without introducing YOLOv8 and feature fusion attention mechanism was 86.94%. When YOLOv8 and feature fusion attention mechanism were introduced in Layer 4 of TSM, the accuracy was increased to 93.75%, indicating that the fusion of high-level features effectively improved the classification performance of the model. And TSM implements the fusion strategy at Layer2 and Layer3 with an accuracy of 93.20%; it implements the fusion strategy at Layer3 and Layer4, such as Figure 6As shown in the figure, the accuracy rate is improved to 94.44%, which further verifies that the deeper the feature fusion is, the more helpful it is to improve the classification of violent behavior detection, reflecting the superiority of the feature fusion strategy.
[0110] In the application of violent behavior detection in actual campus monitoring scenarios, this paper uses two technical solutions: dynamic analysis of motion trajectories and overall understanding of video content, and finally combines the user interface developed based on Streamlit to achieve accurate monitoring and efficient management of campus safety. The implementation method of the system covers usage methods, implementation equipment, and scope of application, etc., forming a complete set of intelligent monitoring solutions.
[0111] The system first processes and analyzes surveillance videos through four main steps: data input, feature extraction, analysis and detection, and result output. In the motion trajectory dynamic analysis solution, monitoring equipment (such as fixed cameras) collects video data in campus scenes in real time and transmits it to the analysis platform. The system uses the YOLOv8 model to detect pedestrian targets in video frames, and uses the BoT-SORT algorithm to track the identity of pedestrians and generate motion trajectory data for each pedestrian. Subsequently, the trajectory data is input into the LSTM model for time series feature extraction to analyze whether there is violent behavior. In addition, bidirectional time series analysis further optimizes the trajectory features and significantly improves the accuracy of detection. Users can intuitively view the processing process through the interface, including detection boxes, trajectory lines, and classification results.
[0112] In the overall video content understanding solution, the system adopts a combination of YOLOv8 and TSM network architecture. YOLOv8 is used to extract 2D spatial features of video frames, while TSM captures spatiotemporal features between frames. Through multi-layer feature fusion and attention mechanism, the system can accurately analyze behaviors in complex scenes. Once violent behavior is detected, the system will mark the relevant areas in the user interface in real time and trigger an alarm. At the same time, the complete analysis results are displayed through classification probability maps and event logs, providing detailed information for campus safety management.
[0113] The operation of the system depends on efficient hardware and software collaboration. Data collection equipment includes high-definition surveillance cameras deployed on campus to ensure the stability and clarity of video collection. The analysis platform uses high-performance GPU servers or edge computing devices to meet the computing needs of deep learning models and provide real-time detection feedback. The user interaction interface is developed by Streamlit, which provides convenient operation experience and functional support, including video uploading, real-time monitoring, detection result display and report export. The back-end logic is implemented through Python, and the interface is seamlessly connected with the deep learning algorithm to ensure the smooth operation of the system.
[0114] In summary, the present invention provides two innovative technical solutions for campus safety management based on the two perspectives of dynamic analysis of motion trajectories and overall understanding of video content by constructing a comprehensive detection system that can integrate multiple features, which can meet the detection of violent behavior in campus surveillance videos, thereby achieving early warning and timely response to violent behavior on campus. The present invention designs a method that integrates target detection and time series features, taking into account detection accuracy, computational complexity, and real-time performance, while making full use of customized data sets for campus scenes, which can meet the actual needs of campus safety management for violent behavior detection.
[0115] The present invention adopts TSM, which is an efficient video understanding model based on time dimension feature offset, which can effectively capture the timing information in the video sequence, thereby enhancing the perception of dynamic behavior. TSM offsets the features of some channels in the time dimension without increasing the computational complexity through the time offset operation, thereby achieving effective information propagation in the time domain. This design enables the model to fully explore the temporal dependencies in the video sequence and makes the flow of temporal information more efficient, thereby achieving good temporal feature extraction effects without requiring a lot of calculations.
[0116] In order to meet the actual needs of campus violence detection algorithms, the present invention combines multiple public data sets with self-built campus scene data sets, and screens and supplements video clips suitable for campus environments in a targeted manner. Finally, a high-quality campus violence data set is designed and constructed, and the data is processed in a refined manner to ensure data diversity and high relevance to practical applications.
[0117] The YOLOv8 used in the present invention as the target detection model has achieved a good balance in detection accuracy, resource utilization and real-time performance, and is suitable for detecting violent behavior in campus monitoring environments. The present invention introduces channel attention and spatial attention mechanisms into the model, allowing the model to focus more on key features in the video frame.
[0118] The present invention adopts the BoT-SORT algorithm, which combines IoU matching with ReID (re-identification) feature fusion to ensure that the motion trajectory of each target can be continuously and accurately tracked in the case of multiple targets. The campus surveillance video violence detection algorithm proposed by the present invention, which is based on the integration of target detection, target tracking and time series analysis, significantly improves the detection accuracy by deeply analyzing the motion trajectory of pedestrians in the video.
[0119] The present invention solves the technical problems existing in the prior art, such as missed detection and false detection caused by insufficient feature utilization, difficulty in balancing real-time performance and accuracy, and poor adaptability in campus monitoring scenarios.
[0120] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting campus violence, characterized in that: The method comprises: S1. Build and combine public data sets and self-built campus scene data sets, filter out campus environment video clips, and build campus scene violence behavior data sets based on campus environment video clips; S2. Detecting violent behavior in campus surveillance videos based on motion trajectory analysis; wherein, constructing and using the YOLOv8 model as a detector to detect pedestrian targets in surveillance videos from the campus scene violent behavior dataset; combining the tracking algorithm to track the pedestrian targets in real time and extract motion trajectories; performing behavioral analysis on the pedestrian targets, analyzing the tracked motion trajectories through the long short-term memory model, extracting abnormal behavior pattern features, and judging the behavior of the pedestrian targets based on the obtained behavior for early warning and response of violent behavior; S3. Set and use the video understanding model TSM to perform video understanding operations, capture the temporal information of the video sequence in the campus scene violence data set, perform time offset operations on channel features in a preset time dimension, and perform local time information exchange in the input feature map of each time step to extract strong temporal dependency features and perceive the dynamic behavior of the pedestrian target; perform feature fusion operations on the output data of the video understanding model TSM and the YOLOv8 model to obtain a fused feature map, which is used to detect campus surveillance line video violence.
2. A campus violence detection method according to claim 1, characterized in that: The S1 includes: S11, screening out violent videos, non-violent videos, and campus scene videos to construct the public data set; S12, using a preset acquisition device to collect surveillance videos of violent behavior to construct the self-built campus scene data set; S13. Refining the campus scene violence behavior dataset.
3. A campus violence detection method according to claim 2, characterized in that: The S13 includes: S131, dividing the videos in the campus scene violence behavior dataset to convert them into frame sequences for time series analysis; S132, using a preset annotation tool to annotate the frame sequence frame by frame to obtain an annotated frame sequence; S133: Perform enhancement processing on the marked frame sequence.
4. A campus violence detection method according to claim 1, characterized in that: In S2, the BoT-SORT algorithm is used to perform IoU matching operations and ReID feature fusion operations, and the motion trajectory of each pedestrian target is tracked in a multi-target situation; In the BoT-SORT algorithm, all the pedestrian targets in the video frame are detected by the YOLOv8 model; According to the IoU matching mechanism of the pedestrian target, the target is tracked between consecutive frames, and the ReID feature vector is obtained and combined to perform an ID assignment operation for each pedestrian target, so as to continuously track the motion trajectory of the pedestrian target in the video stream; In the process of violent behavior identification, motion trajectory data fusion is performed to analyze and obtain behavioral feature difference data and perform feature extraction operations; The feature extraction operation includes: adjusting the aspect ratio of the detection rectangular frame; The trajectory information of all IDs in the original video and the aspect ratio of the detection rectangle of the ID are integrated into comprehensive data.
5. A campus violence detection method according to claim 1, characterized in that: In S2, the YOLOv8 model includes: a channel attention module and a spatial attention module; which are used to focus on processing key features of video frames in the campus scene violence data set. Using the channel attention module, a channel weight vector is generated by global average pooling to adjust the attention of the YOLOv8 model to different feature channels; Using the spatial attention module, a spatial weight map is generated by combining maximum pooling with convolution operations to optimize the YOLOv8 model's attention to different locations in the video frame and focus on areas containing violent behavior; The YOLOv8 model is trained using a frame-level annotation dataset.
6. A campus violence detection method according to claim 1, characterized in that: In S2, the BoT-SORT algorithm is used to track the target, perform IoU matching operations and re-identification (ReID) feature fusion operations, and track the motion trajectory of each pedestrian target in a multi-target situation; In the BoT-SORT algorithm, all the pedestrian targets in the video frame are detected by the YOLOv8 model; According to the IoU matching mechanism of the pedestrian target, the target is tracked between consecutive frames, and the ReID feature vector is obtained and combined to perform an ID assignment operation for each pedestrian target, so as to continuously track the motion trajectory of the pedestrian target in the video stream; In the process of violent behavior identification, motion trajectory data fusion is performed to analyze and obtain behavior feature difference data, and perform feature extraction operations, wherein the feature extraction operation includes: adjusting the aspect ratio of the detection rectangular frame; integrating the trajectory information of all IDs in the original video and the aspect ratio of the detection rectangular frame of the ID into comprehensive data.
7. A campus violence detection method according to claim 1, characterized in that: In S2, a long short-term memory network model LSTM is constructed, wherein the long short-term memory network model LSTM includes: an LSTM layer, a Bi-LSTM layer, an adaptive average pooling layer and a fully connected layer.
8. A campus violence detection method according to claim 7, characterized in that: Set input tensor parameters; set and use the LSTM layer to obtain and capture the time dependency in the sequence data based on the hidden state and the cell state to obtain LSTM output data; Setting and utilizing the Bi-LSTM layer to simultaneously process the forward time series and the backward time series to obtain Bi-LSTM output data; The LSTM output data and the Bi-LSTM output data are concatenated to obtain a forward and bidirectional time series information feature vector, and the last time step hidden state is extracted based on the feature vector; Setting and utilizing the adaptive average pooling layer to perform a pooling operation on the target motion features in the last time step hidden state to obtain a pooling feature vector; Setting and utilizing the fully connected layer, processing the pooled feature vector, integrating the time features extracted by the LSTM layer and the Bi-LSTM layer, and obtaining the fully connected layer output data; The sigmoid activation function is used to map the output data of the fully connected layer to a preset interval, and the probability score of violent behavior is predicted.
9. A campus violence detection method according to claim 1, characterized in that: The feature fusion operation includes: S31, concatenating the 2D features output by the FPN layer P2 of the YOLOv8 model and the 3D features output by the backbone layer layer2 of the time-frequency understanding model TSM in a preset channel dimension to obtain tensor concatenated features; S32, passing the tensor concatenation features through two convolutional layers to obtain a feature tensor F1; S33, reshaping the feature tensor F1 into a first three-dimensional feature FQ for matrix operation; S34, performing a transposition operation on the first three-dimensional feature FQ to obtain a transposed feature FK; S35, performing softmax normalization on the dot product of the first three-dimensional feature FQ and the transposed feature FK to generate an attention map Fat; S36, reshape the feature tensor F1 to obtain a first three-dimensional feature FV, and multiply it by the attention map Fat to obtain the fused feature map; S37, adding the fused feature map to the feature tensor F1; after two convolution operations, completing the feature dimension recovery operation and absorbing the effective information of the 2D feature.
10. A campus violence detection system, characterized in that: The system comprises: The violent behavior dataset construction module is used to construct and combine public datasets and self-built campus scene datasets, filter out campus environment video clips, and construct campus scene violent behavior datasets based on campus environment video clips; A motion trajectory analysis module is used to detect violent behavior in campus surveillance videos based on motion trajectory analysis; wherein, a YOLOv8 model is constructed and used as a detector to detect pedestrian targets in surveillance videos from the campus scene violent behavior dataset; the pedestrian targets are tracked in real time in combination with a tracking algorithm to extract motion trajectories; behavioral analysis is performed on the pedestrian targets, and the motion trajectories obtained by tracking are analyzed by a long short-term memory model to extract abnormal behavior pattern features, and the behavior of the pedestrian targets is judged accordingly for violent behavior warning and response, and the motion trajectory analysis module is connected to the violent behavior dataset construction module; The video understanding module is used to set and use the video understanding model TSM to perform video understanding operations, capture the temporal information of the video sequence in the campus scene violent behavior dataset, perform time offset operations on channel features in a preset time dimension, and perform local time information exchange in the input feature map of each time step to extract strong temporal dependency features and perceive the dynamic behavior of the pedestrian target; perform feature fusion operations on the output data of the video understanding model TSM and the YOLOv8 model to obtain a fused feature map, and use it to detect violent behavior in campus surveillance line videos. The video understanding module is connected to the violent behavior dataset construction module.
Citation Information
Patent Citations
Violence detection method and device based on neural network, equipment and medium
CN112287754A
Distributed monitoring-based violent behavior detection system and detection method
CN113269076A
Cited By
Campus safety behavior management system based on visual identification
CN122454126A