Method for detecting abnormal behaviors of personnel around rivers based on deep learning image processing
Through self-supervised algorithm based on deep learning and multimodal feature extraction network, the behavior vector library and distance discrimination set are constructed, which solves the problem of unrecognized unknown behavior types in the prior art, and realizes high-accuracy abnormal behavior detection around rivers.
Patent Information
- Application Number
- CN202510376633.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The prior art cannot effectively identify unknown behavior types that are not in the training set, resulting in low accuracy of abnormal behavior detection around rivers.
Using deep learning-based image processing method, a self-supervised algorithm and multimodal feature extraction network is used to construct a behavior vector library and a distance discrimination set to achieve judgment of all behavior types, regardless of whether they belong to a known category.
It enhances the generalization ability and adaptability of the model, can effectively identify unknown behavior types, and improves the accuracy of abnormal behavior detection around the river.
Smart Images

Figure CN119888868B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graphic data reading and processing, and particularly to a method for detecting abnormal behaviors of personnel around rivers based on deep learning image processing. Background Art
[0002] With the acceleration of the urbanization process, the safety issues in river areas have become increasingly prominent. As an important part of public safety, abnormal behaviors around rivers (such as climbing guardrails, falling into the river, etc.) not only pose a threat to human lives but may also trigger secondary disasters or accidents. Traditional river safety management methods usually rely on manual patrols and real-time observations by monitoring personnel. However, such methods are inefficient, limited by human resources, and prone to missed reports due to negligence, making it difficult to meet the requirements of modern safety management. With the breakthrough of deep learning technology in the field of computer vision, automated abnormal behavior detection methods based on deep learning have begun to become the key technologies to solve this problem.
[0003] In recent years, the rapid development of deep learning technology, especially the breakthrough in the field of computer vision, has provided new technical support for security monitoring. Deep learning models based on object detection, with their efficient end-to-end detection capabilities, have been widely applied in multiple fields such as traffic monitoring and security systems. They can achieve high-precision object recognition and positioning in complex backgrounds, providing a technical basis for real-time detection of abnormal behaviors around rivers. However, the complexity of river scenes (such as light changes, frequent occlusions, water surface reflections, etc.) poses higher requirements for the robustness of detection models. Relying solely on single-frame object detection is difficult to accurately identify dynamic behavior features and potential abnormal patterns. Therefore, traditional object detection methods face challenges in adapting to complex scenes and dynamic changes. Although there are currently some abnormal behavior recognition methods based on object detection, they usually rely on predefined behavior categories and cannot effectively identify unknown behavior types not in the training set, resulting in low detection accuracy. Therefore, a method for detecting abnormal behaviors of personnel around rivers based on deep learning image processing is developed to solve the above problems. Summary of the Invention
[0004] The present invention proposes a method for detecting abnormal behaviors of personnel around rivers based on deep learning image processing to solve the problem that the existing technology cannot effectively identify unknown behavior types not in the training set, resulting in low detection accuracy.
[0005] The present invention achieves the above object through the following technical solutions:
[0006] A method for detecting abnormal behaviors of personnel around rivers based on deep learning image processing according to the present invention includes:
[0007] Obtain the first information and the second information, where the first information is a training data set with labels, the training data set includes historical monitoring image data around the river, the labels include normal behavior labels and abnormal behavior labels, and the second information includes real-time monitoring video data around the river;
[0008] Preprocess the first information to obtain preprocessed data;
[0009] Train a preset self-supervised key point extraction network according to the preprocessed data, including inputting the preprocessed data into a preset object detection model to output the person object features in the surrounding area, training the object detection model based on the intersection over union loss function, inputting the preprocessed data into a preset key point detection model to output the person dynamic features, training the key point detection model based on the key point coordinate loss function, using the person object features and the person dynamic features as input samples to input into a preset multi-modal self-supervised learning network to generate the embedding vector of each data, and training the multi-modal self-supervised learning network based on the contrastive learning method;
[0010] Input the preprocessed data into the multi-modal self-supervised learning network to output a behavior vector library, classify and label the embedding vectors in the behavior vector library according to the labels, calculate the distances between the embedding vectors corresponding to each normal behavior and the embedding vectors corresponding to each abnormal behavior, and sort all the distances from small to large to obtain a distance discrimination set;
[0011] Preprocess the second information and then input it into the trained self-supervised key point extraction network to output the corresponding real-time embedding vector, calculate the distance between the real-time embedding vector and the embedding vector corresponding to the abnormal behavior, calculate the sorting ratio of the distance in the distance discrimination set, compare the sorting ratio with a preset alarm threshold, and if the sorting ratio is less than or equal to the alarm threshold, then give an alarm for abnormal behavior.
[0012] Further, obtaining the first information includes:
[0013] Obtain the video data continuously recorded by 2 high-definition infrared monitoring cameras for 72 hours. The resolution of the 2 high-definition infrared monitoring cameras is 2560×1440, the viewing angle is 110°, the installation height is 3.5 meters, and the inclination angle is 15°;
[0014] Extract the key frames from the video data based on the FFmpeg program module, extract 2 frames of images per second, and generate a standardized image data set of 640×640 pixels;
[0015] Gaussian filtering and denoising are performed on the standardized image data set, and the target is annotated using the LabelImg tool to generate an XML annotation file to obtain the training data set with labels.
[0016] Furthermore, the pre-processing step comprises:
[0017] Performing image correlation analysis on initial image data to remove low-quality data to obtain a key frame image sequence, wherein the initial image data is the first information or the second information;
[0018] The Laplace operator variance method is used to calculate the clarity score of each image in the key frame image sequence, and the images whose clarity scores are lower than a preset threshold are marked as low-quality frames and removed to obtain clear image data;
[0019] Using a contrast-limited adaptive histogram equalization method to optimize each frame of the clear image data to obtain an optimized image;
[0020] The optimized image is processed based on an illumination normalization model of a convolutional neural network.
[0021] Further, according to the pre-processed data, the preset target detection model is input, the target features of the people in the surrounding area are output, and the target detection model is trained based on the intersection-over-union loss function, including:
[0022] The preprocessed data is divided into a training set and a validation set in a ratio of 80%:20%;
[0023] Construct a target detection model, which includes a resnet50 backbone network, an FPN feature pyramid network and a detection head. The target detection model is used to first extract features from the input data through the resnet50 backbone network, then perform multi-scale feature fusion on the extracted features through the FPN feature pyramid network, and finally input the fused multi-scale features into the detection head to perform target positioning and output the detection results;
[0024] According to the training set, the constructed target detection model is trained to a specified round based on the intersection-over-union loss function and the data enhancement strategy, and the intersection-over-union loss function is:
[0025] ;
[0026] ;
[0027] in, and are the predicted bounding box and the true bounding box respectively, N is the number of detected targets, is the intersection-and-union ratio, is the weight factor, Area of Overlap represents the area of the overlapping region between the predicted bounding box and the ground truth bounding box, Area of Union represents the total area of the predicted bounding box and the ground truth bounding box, and the value range of IoU is [0, 1].
[0028] Further, the key point coordinate loss function is as follows:
[0029] ;
[0030] wherein, is the detected human key point coordinates, is the labeled key point coordinates.
[0031] Further, taking the person target feature and the person dynamic feature as input samples and inputting them into a preset multi-modal self-supervised learning network to generate an embedding vector for each data, and training the multi-modal self-supervised learning network based on the contrast learning method, including:
[0032] Construct a multi-modal self-supervised learning network, which includes an image feature extraction module, a key point coordinate extraction module, a feature fusion module, and an embedding vector module. The image feature extraction module is used to extract image feature vectors from the person target feature, the key point coordinate extraction module is used to extract key point feature vectors from the person dynamic feature, the feature fusion module is used to fuse the image feature vectors and key point feature vectors, and the embedding vector module is used to perform feature embedding on the fused vectors. The embedding vector module includes two fully connected layers;
[0033] Perform multiple different augmentations on the input samples based on the data augmentation method, including rotation, translation, flipping, and random occlusion;
[0034] Perform multiple different augmentations on each input sample to generate positive sample pairs, and all other samples in the training set form negative samples with them, and use the contrast loss function to optimize the multi-modal self-supervised learning network;
[0035] The contrast loss function is as follows:
[0036] ;
[0037] wherein, and are the embedding vectors of the positive sample pair respectively, is the embedding vector of the negative sample, sim represents the cosine similarity between the embedding vectors, is the temperature hyperparameter, which is used to control the sensitivity of the distance;
[0038] During the training process, when the contrast loss function converges, the multi-modal self-supervised learning network is trained and completed.
[0039] Further, calculate the distance between the embedding vectors corresponding to each normal behavior and the embedding vectors corresponding to each abnormal behavior The calculation formula is as follows:
[0040] ;
[0041] ;
[0042] represents the embedding vector of the normal behavior, represents the embedding vector of the abnormal behavior, j represents the index of the abnormal behavior embedding vector, and M represents the total number of abnormal behavior embedding vectors.
[0043] The beneficial effects of the present invention are as follows:
[0044] The method for detecting abnormal behaviors of personnel around rivers based on deep learning image processing proposed by the present invention adopts a self-supervised algorithm, combines a multi-modal feature extraction network, and analyzes personnel images and dynamic behavior characteristics through feature embedding. Through the behavior vector library and distance discrimination set constructed during the training process, this method realizes the judgment of all behavior types, regardless of whether they belong to known categories. Through self-supervised learning, the behavior category is judged through the behavior embedding vector, rather than the known behaviors marked in the training set, thereby enhancing the generalization ability and adaptability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flowchart of a method for detecting abnormal behaviors of personnel around rivers based on deep learning image processing according to the present invention.
[0046] Figure 2 is the algorithm flow of step 3 in the embodiment of the present invention.
[0047] Figure 3 is the self-supervised model structure of step 3 in the embodiment of the present invention.
[0048] Figure 4 is the algorithm flow of step 5 in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein can generally be arranged and designed in a variety of different configurations.
[0050] Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0051] It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0052] In addition, terms such as "first", "second", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance.
[0053] The following will describe in detail the specific embodiments of the present invention in conjunction with the accompanying drawings.
[0054] As Figure 1 shown, a flowchart of a method for detecting abnormal behaviors of people around rivers based on deep learning image processing according to the present invention specifically includes:
[0055] S1: Obtain first information and second information, where the first information is a training data set with labels, the training data set includes historical monitoring image data around the river, the labels include normal behavior labels and abnormal behavior labels, and the second information includes real-time monitoring video data around the river;
[0056] S2: Preprocess the first information to obtain preprocessed data;
[0057] S3: Train a preset self-supervised key point extraction network according to the preprocessed data, including inputting the preprocessed data into a preset object detection model to output the person target features in the surrounding area, training the object detection model based on the intersection over union loss function, inputting the preprocessed data into a preset key point detection model to output the person dynamic features, training the key point detection model based on the key point coordinate loss function, using the person target features and the person dynamic features as input samples to input into a preset multi-modal self-supervised learning network to generate an embedding vector for each data, and training the multi-modal self-supervised learning network based on the contrastive learning method;
[0058] S4: Input the preprocessed data into the multi-modal self-supervised learning network to output a behavior vector library, classify and label the embedding vectors in the behavior vector library according to the labels, calculate the distances between the embedding vectors corresponding to each normal behavior and the embedding vectors corresponding to each abnormal behavior, and sort all the distances from small to large to obtain a distance discrimination set;
[0059] S5: Preprocess the second information and input it into the trained self-supervised key point extraction network to output the corresponding real-time embedding vector. Calculate the distance between the real-time embedding vector and the embedding vector corresponding to the abnormal behavior, calculate the sorting ratio of the distance in the distance discrimination set, and compare the sorting ratio with a preset alarm threshold. If the sorting ratio is less than or equal to the alarm threshold, an abnormal behavior alarm is triggered.
[0060] The specific implementation steps of the present invention are as follows:
[0061] Step 1: Select 2 Hikvision DS-2CD2T47G1 high-definition cameras with a resolution of 2560×1440, supporting night infrared monitoring and a viewing angle of 110°. Each camera is equipped with a waterproof housing and a shockproof bracket to ensure adaptation to the long-term outdoor environment. The camera installation height is 3.5 meters, and the inclination angle is set to 15°. Use a laser rangefinder to ensure that the camera coverage has no dead angles. Conduct continuous video recording for 72 hours, with the video format being MP4 and the frame rate being 25 frames per second, stored on a 2TB SSD hard drive. Use FFmpeg to extract the key frames in the video, extract 2 frames of images per second, and generate a standardized image dataset of 640×640 pixels. Perform Gaussian filtering on the images to remove noise, and use the LabelImg tool for object annotation to generate XML annotation files. Finally, generate approximately 100,000 frame image data to provide a basis for subsequent model training.
[0062] Step 2: Perform preprocessing operations on the collected original video data to improve data quality and convert the data into a standard format suitable for model training. Specifically include:
[0063] Step 2.1: Remove low-quality data based on the correlation of the collected images. By analyzing the sequence of key frame images collected, first use an image sharpness evaluation algorithm to score the sharpness of each frame. Use the variance method of the Laplacian operator to calculate the sharpness score of the image . Images with sharpness lower than the preset threshold ( ) will be marked as low-quality frames and removed. And use an image optimization algorithm based on comparative analysis to improve the image quality. Use the Contrast Limited Adaptive Histogram Equalization (CLAHE) method to optimize each frame of the image, enhancing the image contrast while avoiding detail loss caused by over-enhancement.
[0064] Step 2.2: Construct a de-illumination normalization neural network to eliminate the interference of illumination changes. For the problem of non-uniform illumination caused by strong light or shadows in some images, use a convolutional neural network (CNN)-based illumination normalization model, and use existing publicly available trained algorithms to process the data, enhancing the contrast and sharpness of the target area.
[0065] Step 3: As Figure 2 shown, use the self-supervised key point extraction network to locate the people in the image and embed the behavioral features, and train the object detection model. Specifically, it includes:
[0066] Step 3.1: Based on the image dataset and annotation files generated in the previous two steps, divide the data into a training set and a validation set in a ratio of 80%:20%. The object detection network consists of a resnet50 backbone network, an FPN feature pyramid network, and a detection head. For key point detection, use the trained human key point extraction to obtain the skeletal key points of the target area. After that, obtain the behavioral vector embedding through a multi-modal feature extraction network.
[0067] During the training process, the loss function of the object detection algorithm is as follows:
[0068] Bounding box loss:
[0069] ;
[0070] Where and are the predicted bounding box and the ground truth bounding box respectively; N is the number of detected objects; is the weight factor; used to assign different weights to different objects; is the Intersection over Union (IoU), and its calculation formula is as follows:
[0071] ;
[0072] Where Area of Overlap represents the area of the overlapping region between the predicted bounding box and the ground truth bounding box, Area of Union represents the total area of the predicted bounding box and the ground truth bounding box (excluding the overlapping part), and the value range of IoU is [0, 1]. The closer the value is to 1, the more similar the two boxes are.
[0073] During the training process, add data augmentation strategies (such as rotation, cropping, brightness adjustment, and random occlusion simulation) to improve the generalization ability of the model for complex scenarios.
[0074] Use the training data to train to the specified number of epochs to make the model reach the convergence state and obtain the trained object detection algorithm model.
[0075] The loss function of the key point detection model is:
[0076] ;
[0077] Where Are the coordinates of the detected human key points (head, both hands, center point, both feet). Are the coordinates of the labeled key points.
[0078] Step 3.2: As Figure 3 shown, considering that the data obtained by the above model includes both personnel image data and key point coordinate data, it is difficult for a single model to process these two data types. Therefore, the present invention constructs a multi-modal self-supervised learning network for feature extraction and embedding. The structure of this network is as follows:
[0079] Image feature extraction module: ResNet18, input data of 224*224*3, output features of 64 dimensions ;
[0080] Key point coordinate extraction module: Input a key point feature vector of length 12, and after passing through two fully connected layers (the first layer outputs 32 dimensions, using ReLU activation; the second layer outputs 64 dimensions), obtain a 64-dimensional feature vector ;
[0081] Feature fusion module: Concatenate the image feature vector and the key point feature vector to obtain a 128-dimensional fused feature vector .
[0082] Embedding vector module: Input a 128-dimensional fused feature vector , and obtain the final embedding vector e through two fully connected layers (the first layer outputs 128 dimensions, using ReLU activation; the second layer outputs 256 dimensions) for representing the feature embedding of the input image and key point data.
[0083] In order to construct an effective self-supervised learning task, the present invention adopts data augmentation methods, including rotation, translation, flipping and random occlusion, to perform multiple different augmentations on each input sample to generate positive sample pairs. All the remaining samples are used as negative samples, and the contrastive loss (ContrastiveLoss) is used to optimize the embedding space. The specific contrastive loss function is as follows:
[0084]
[0085] Among them, and are the embedding vectors of the positive sample pair respectively, is the embedding vector of the negative sample, sim represents the cosine similarity between the embedding vectors, is the temperature hyperparameter used to control the sensitivity of the distance.
[0086] During the training process, when the loss function reaches convergence, that is, when the model no longer decreases further, it indicates that the training is completed. At this time, use the trained model to process the data obtained in step 4.1 again, and save all the generated embedding vectors. At the same time, label the embedding vectors with the labels in step 1, and finally obtain the behavior vector library.
[0087] Step 3.3: Calculate the L2 distance between the embedding vectors of each normal behavior and the embedding vectors of all abnormal behaviors, and take the average of the L2 distances of all abnormal behaviors. The specific steps are as follows:
[0088] Let the combined behavior vector library of normal behaviors be and the combined behavior vector library of abnormal behaviors be .
[0089] For each embedding vector of a normal behavior , calculate its L2 distance from the embedding vectors of all abnormal behaviors :
[0090] ;
[0091] where represents the L2 norm between two embedding vectors, that is, the Euclidean distance.
[0092] For each embedding vector of a normal behavior, calculate its average L2 distance to the embedding vectors of all abnormal behaviors:
[0093] ;
[0094] After calculating the average L2 distance between the embedding vectors of all normal behaviors and the embedding vectors of abnormal behaviors, sort these average distances in ascending order to obtain the average distance ranking set of normal behaviors to abnormal behaviors:
[0095] ;
[0096] Finally, obtain the result of the average L2 distance of each normal behavior embedding vector to all abnormal behavior embedding vectors arranged in ascending order as the distance discrimination set.
[0097] Step 4: Deploy the optimized object detection and human keypoint detection models to the monitoring device, integrate them with the real-time video stream, and improve the detection and analysis efficiency through hardware acceleration or edge computing to ensure real-time response. After the training of the optimized object detection model and the human keypoint detection module is completed, they are exported in the ONNX (Open Neural Network Exchange) format for efficient operation on hardware devices. During the deployment process, the model is loaded onto the NVIDIA Jetson Xavier NX edge computing device, which is equipped with a 6-core CPU, 384-core GPU, and 8GB of memory, capable of meeting the high-performance requirements for real-time inference and keypoint analysis. The present invention integrates a real-time video stream processing module, and the video stream is transmitted through the RTSP protocol and decoded using OpenCV.
[0098] Step 5: As Figure 4 shown, use the deployed model to detect, analyze the behavior, and perform behavior embedding on the personnel in the real-time monitoring screen, generate behavior feature data, and compare it with the behavior vector library in the database to determine whether the behavior is abnormal. The deployed object detection model receives the real-time monitoring video stream, and each frame of the video is subjected to personnel target detection through the object detection model to generate the bounding box coordinates of the target, where and are the center point coordinates of the target bounding box, and w and h are the width and height of the bounding box. At the same time, the human keypoint detection module extracts the keypoint coordinates , where each represents the x and y coordinates of the keypoint. Subsequently, the personnel image and the human keypoint coordinates are input into the self-supervised learning network designed by the invention to generate a real-time embedding vector , and calculate the L2 distance between it and the embedding vectors of all abnormal behaviors in the behavior vector library obtained in Step 3.2. The specific L2 distance calculation formula is:
[0099] ;
[0100] where, is the real-time embedding vector, and is the embedding vector of the abnormal behavior.
[0101] Calculate the said distance in the said distance discrimination set , compare the sorting ratio of the said distance with the preset alarm threshold θ, and if the sorting ratio is less than or equal to the alarm threshold, the behavior is determined to be an abnormal behavior and an alarm is triggered.
[0102] Step 6: All relevant data determined to be abnormal behaviors are saved in real time, including the personnel location , Personnel Image , Behavioral Feature Embedding Vector , Key Point Coordinates , Detection Timestamp , and Abnormality Judgment Result of the Behavior (where indicates abnormal behavior, indicates normal behavior) Conduct a detailed review of the saved data based on the abnormality judgment result and the real-time monitored video, and mark whether it actually belongs to abnormal behavior. The marked result will update the annotation in the original dataset and be saved together with the original data as a new dataset. Each annotation result will include detailed review opinions, corrected behavior categories, and relevant time information to ensure the accuracy and integrity of the data. Finally, the corrected and annotated results will form a new dataset , and this dataset will be used for subsequent model training, validation, optimization, and continuous improvement of performance.
[0103] Step 7: Regularly summarize and analyze the detection results to identify and summarize the time and location distributions of abnormal behaviors. For example, by analyzing the time periods and location areas where abnormal behaviors occur, identify which time periods and location areas frequently have abnormal behaviors. This information will be used to troubleshoot the influence of non-human factors, such as environmental light changes, equipment failures, or interference from other external factors, thereby optimizing the stability and accuracy of the self-supervised key point extraction network.
[0104] At the same time, when the number of new datasets reaches a certain level, use these data to fine-tune the multi-modal self-supervised learning network to further improve the accuracy and adaptability of the model. Specifically, the new dataset will be input into the self-supervised network for update and training, re-perform the embedding vector generation process in Step 3.2, and recalculate the behavior vector library and abnormal behavior embedding vector based on the new data. Then, update the distance discrimination set in Step 3.3, and optimize the detection accuracy and generalization ability of abnormal behaviors by sorting and adjusting the threshold. Through this periodic fine-tuning and update process, the self-supervised key point extraction network will continuously improve its abnormal behavior detection ability, ensure that it can adapt to changes in different environments and behavior patterns, and thus enhance the recognition accuracy of unknown or unseen abnormal behaviors.
[0105] In summary, the present invention realizes the accurate detection of personnel targets in the areas around rivers, and identifies abnormal behaviors through the dynamic analysis of skeletal key points and the calculation of behavior embedding vectors. By using the object detection model and the behavior embedding vectors generated by self-supervised learning, the present invention can perform real-time abnormal behavior detection and alarm in complex scenarios around rivers. Through the combination of object detection and behavior embedding, not only can the target be accurately located, but also the motion patterns, behavior characteristics and their dynamic changes of the target can be captured in real time, effectively identifying abnormal behaviors. This technology shows high robustness in the face of complex environments such as light changes, target occlusion and fast movement, meeting the dual requirements of real-time performance and accuracy. Through multiple optimization strategies such as key frame extraction, image preprocessing and key point sequence analysis, the data quality and model adaptability are further improved. Combining the self-supervised learning method of behavior embedding, it can continuously adapt to new abnormal behavior patterns by learning the embedding vectors of normal behaviors, ensuring the stability and accuracy of the self-supervised key point extraction network in practical applications. This technology provides efficient and stable technical support for abnormal behavior recognition, and can reliably perform real-time alarm and data recording in complex environments.
[0106] The present invention proposes a complete technical solution for the safety management around rivers, comprehensively applying technologies such as software and hardware co-design, real-time video stream processing, object detection, dynamic behavior analysis and behavior embedding vectors generated by self-supervised learning, significantly improving the stability and efficiency of the self-supervised key point extraction network. Combining the intelligent recognition and automatic alarm mechanism of the dynamic characteristics of human key points, the self-supervised key point extraction network can efficiently adapt to the requirements of different scenarios. It has important value in reducing potential safety hazards and improving the level of public safety management, and can cope with common complex factors (such as light changes, the influence of obstacles, etc.) in the environment around rivers. The present invention not only improves the intelligent and refined management capabilities of the monitoring system, but also can continuously expand the ability to identify new abnormal behaviors, ensuring the comprehensiveness and accuracy of public safety management. This technical solution is widely applicable to fields such as river safety monitoring and public safety protection, providing strong technical support for realizing intelligent and refined safety management.
[0107] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for detecting abnormal behavior of people around rivers based on deep learning image processing, characterized in that: include: Acquire first information and second information, wherein the first information is a training data set with labels, the training data set includes historical monitoring image data around the river, the labels include normal behavior labels and abnormal behavior labels, and the second information includes real-time monitoring video data around the river; Preprocessing the first information to obtain preprocessed data; Training a preset self-supervised key point extraction network according to the preprocessed data, including inputting the preprocessed data into a preset target detection model, outputting the personnel target features of the surrounding area, and training the target detection model based on the intersection-over-union loss function, inputting the preprocessed data into a preset key point detection model, outputting the personnel dynamic features, and training the key point detection model based on the key point coordinate loss function, inputting the personnel target features and the personnel dynamic features as input samples into a preset multimodal self-supervised learning network to generate an embedding vector for each data, and training the multimodal self-supervised learning network based on a contrastive learning method; Input the preprocessed data into the trained self-supervised key point extraction network, output a behavior vector library, classify and label the embedded vectors in the behavior vector library according to the label, calculate the average distance between the embedded vector corresponding to each normal behavior and the embedded vectors corresponding to all abnormal behaviors, and sort all the average distances from small to large to obtain a distance discriminant set; The second information is preprocessed and then input into the trained self-supervised key point extraction network, the corresponding real-time embedding vector is output, the distance between the real-time embedding vector and the embedding vector corresponding to the abnormal behavior in the behavior vector library is calculated, the ranking ratio of the distance in the distance discrimination set is calculated, the ranking ratio is compared with the preset alarm threshold, and if the ranking ratio is less than or equal to the alarm threshold, an abnormal behavior alarm is issued.
2. The method for detecting abnormal behavior of people around a river based on deep learning image processing according to claim 1 is characterized in that: Acquiring the first information includes: Obtain video data recorded continuously for 72 hours by two high-definition infrared surveillance cameras, with a resolution of 2560×1440, a viewing angle of 110°, an installation height of 3.5 meters, and an inclination of 15°; Extract key frames from the video data based on the FFmpeg program module, extract 2 frames of images per second, and generate a standardized image data set of 640×640 pixels; Gaussian filtering and denoising are performed on the standardized image data set, and the target is labeled using the LabelImg tool to generate an XML labeling file to obtain the training data set with labels.
3. The method for detecting abnormal behavior of people around a river based on deep learning image processing according to claim 1 is characterized in that: The pre-processing step comprises: Performing image correlation analysis on initial image data to remove low-quality data to obtain a key frame image sequence, wherein the initial image data is the first information or the second information; The Laplace operator variance method is used to calculate the clarity score of each image in the key frame image sequence, and the images whose clarity scores are lower than a preset threshold are marked as low-quality frames and removed to obtain clear image data; Using a contrast-limited adaptive histogram equalization method to optimize each frame of the clear image data to obtain an optimized image; The optimized image is processed based on an illumination normalization model of a convolutional neural network.
4. The method for detecting abnormal behavior of people around a river based on deep learning image processing according to claim 1 is characterized in that: The preprocessed data is input into a preset target detection model, the target features of people in the surrounding area are output, and the target detection model is trained based on the intersection-over-union loss function, including: The preprocessed data is divided into a training set and a validation set in a ratio of 80%:20%; Construct a target detection model, which includes a resnet50 backbone network, an FPN feature pyramid network and a detection head. The target detection model is used to first extract features from the input data through the resnet50 backbone network, then perform multi-scale feature fusion on the extracted features through the FPN feature pyramid network, and finally input the fused multi-scale features into the detection head to perform target positioning and output the detection results; According to the training set, the constructed target detection model is trained to a specified round based on the intersection-over-union loss function and the data enhancement strategy. for: ; ; in, and are the predicted bounding box and the true bounding box respectively, N is the number of detected targets, is the intersection-and-union ratio, is the weight factor, Area of Overlap represents the overlapping area of the predicted bounding box and the true bounding box, Area of Union represents the total area of the predicted bounding box and the true bounding box, and the IoU value range is [0, 1].
5. The method for detecting abnormal behavior of people around a river based on deep learning image processing according to claim 1 is characterized in that: The key point coordinate loss function for: ; in, is the detected human key point coordinates, The coordinates of the key points of the annotation.
6. The method for detecting abnormal behavior of people around a river based on deep learning image processing according to claim 1 is characterized in that: Inputting the personnel target features and the personnel dynamic features as input samples into a preset multimodal self-supervised learning network to generate an embedding vector for each data, and training the multimodal self-supervised learning network based on a contrastive learning method, including: Constructing a multimodal self-supervised learning network, the multimodal self-supervised learning network includes an image feature extraction module, a key point coordinate extraction module, a feature fusion module, and an embedding vector module, the image feature extraction module is used to extract image feature vectors for the personnel target features, the key point coordinate extraction module is used to extract key point feature vectors for the personnel dynamic features, the feature fusion module is used to fuse the image feature vectors and the key point feature vectors, the embedding vector module is used to embed features for the fused vectors, and the embedding vector module includes two fully connected layers; Based on the data augmentation method, the input samples are enhanced in multiple different ways, including rotation, translation, flipping and random occlusion; Performing multiple different augmentations on each input sample to generate positive sample pairs, with all other samples in the training set constituting negative samples, and optimizing the multimodal self-supervised learning network using a contrastive loss function; Contrastive loss function as follows: ; in, and are the embedding vectors of the positive sample pairs, is the embedding vector of the negative sample, sim represents the cosine similarity between the embedding vectors, is the temperature hyperparameter, which controls the sensitivity of distance; During the training process, when the contrast loss function reaches convergence, the multimodal self-supervised learning network training is completed.
7. The method for detecting abnormal behavior of people around a river based on deep learning image processing according to claim 1 is characterized in that: Calculate the average distance between the embedding vector corresponding to each normal behavior and the embedding vector corresponding to all abnormal behaviors The calculation formula is as follows: ; ; in, Embedding vector representing normal behavior, represents the embedding vector of abnormal behavior, j represents the index of the abnormal behavior embedding vector, and M represents the total number of abnormal behavior embedding vectors.
8. The method for detecting abnormal behavior of people around a river based on deep learning image processing according to claim 1 is characterized in that: It also includes real-time storage of all data determined to be abnormal behavior, including personnel location, personnel image, behavior characteristics and system time, marking whether it is indeed abnormal behavior, and saving the marking results and data as a new data set.
9. The method for detecting abnormal behavior of people around a river based on deep learning image processing according to claim 8 is characterized in that: It also includes fine-tuning the self-supervised key point extraction network using a new dataset.
Citation Information
Patent Citations
Single classifier anomaly detection method based on multilayer random neural network
CN109858509A
Airport abnormal behavior detection system and method based on multi-modal information fusion
CN111460917A