Low-cost anomaly detection method and system based on multiple cameras
The method and system optimize multi-camera anomaly detection in low-speed scenarios by using wide-angle and narrow-angle cameras with a lightweight model and historical scoring to reduce costs and resource waste, ensuring efficient and accurate detection.
Patent Information
- Application Number
- CN202510487862.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Multi-channel camera systems have redundant problems in cost and efficiency optimization, computing resource optimization and invalid detection in low-speed scenarios, making it difficult to achieve efficient and accurate abnormal detection in low-speed abnormal detection scenarios.
The multi-channel low-cost AHD camera layout is adopted, combined with the lightweight YOLOv8n-ADL model and multi-field dual-frame detection strategy, and dynamic switching detection modes through hierarchical detection and historical anomaly scores are optimized to optimize hardware cost and computing resource utilization.
It realizes all-round field of view coverage in low-speed scenarios, reduces hardware costs, ensures efficient and accurate abnormal detection, and reduces invalid detection through mode switching, and improves resource utilization efficiency.
Smart Images

Figure CN120321384A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent robots, and in particular to a low-cost anomaly detection method and system based on multiple cameras. Background Art
[0002] With the rapid development of artificial intelligence and robot technology, robots are increasingly widely used in various industries. Especially in the security field, intelligent patrol robots, with their autonomous perception and decision-making capabilities, have become an important tool for enhancing the security protection level. In security applications, intelligent patrol robots detect anomalies in the external environment in real time by carrying cameras, including abnormal pedestrian behaviors, environmental anomalies, etc. The vision of anomaly detection often focuses on the accurate coverage of the front vision. However, although the traditional single-camera system can provide certain monitoring functions, there are still obvious vision limitations: the narrow-angle camera can clearly capture distant targets, but it is difficult to take into account wide-area perception; the wide-angle camera can cover the surrounding area, but the details of distant targets are severely lost. To make up for the deficiency in the vision range, more and more robots have adopted a multi-camera system, combining multiple narrow-angle cameras and wide-angle cameras to provide comprehensive monitoring at different distances, enabling the robot to clearly detect abnormal situations at close range and capture abnormal targets in the distance, providing more efficient support for security work.
[0003] However, while the anomaly detection system based on multiple cameras provides a wider vision, it also brings many challenges:
[0004] First, the multi-camera system faces challenges in cost and performance optimization in low-speed scenarios:
[0005] In low-speed scenarios, the high-cost FPD-Link / GMSL camera solution adopted by some existing systems has a significant cost-performance imbalance problem. For example, in the prior art, a robot obtains 1280×720@30fps image data output by a GMSL camera through a SerDes interface for target detection. However, in low-speed scenarios, the degree of motion blur is greatly reduced, and mainstream deep learning models can usually maintain detection accuracy with lower resolutions. The video bandwidth provided by the GMSL / FPD-Link solution exceeds the actual demand, resulting in serious idle hardware resources in low-speed scenarios. Moreover, the high procurement cost of the camera directly raises the overall cost of the system, which conflicts with the cost sensitivity generally required in low-speed scenarios. Therefore, how to construct a balanced "resolution-bandwidth-cost" solution suitable for low-speed scenarios through sensor selection and hardware architecture reconstruction has become a key problem that needs to be solved urgently.
[0006] Second, the multi-camera system faces challenges in computational resource optimization in anomaly detection scenarios:
[0007] In the anomaly detection scenario, the resource overload problem caused by the parallel processing of multiple cameras is particularly prominent. Especially for end-side devices with limited computing resources, the real-time detection of multiple cameras will occupy a large amount of GPU resources; and even for end-side devices with rich computing resources, such as Jetson AGX Orin, it is still necessary to ensure that there are sufficient computing resources to ensure the normal operation of other core algorithms, such as navigation and obstacle avoidance algorithms. Therefore, how to ensure the accuracy of the anomaly detection algorithm while lightweight designing the model to ensure support for parallel detection of multiple cameras with low resource consumption has become a key issue that needs to be solved urgently.
[0008] Third, the multi-channel camera system faces the problem of invalid detection redundancy in abnormal detection scenarios:
[0009] In a multi-channel camera system, the probability of abnormal events is extremely low. When the traditional fixed-period multi-camera polling detection mechanism is adopted, the system needs to continuously perform high-frequency invalid calculations on scenes without abnormalities, which will lead to idle waste of computing resources and reduced computing efficiency. This redundant detection mechanism not only increases the resource burden of the system, but also limits the coordinated analysis capabilities of multiple cameras. Therefore, how to reduce the detection of scenes without abnormalities by each camera has become a problem that needs to be solved urgently.
[0010] Therefore, there is an urgent need to develop a low-cost anomaly detection method and system based on multi-channel cameras to meet the needs of low-speed anomaly detection scenarios and improve the overall performance of the system. Summary of the invention
[0011] In order to solve the technical problems existing in the above-mentioned prior art, the present invention proposes a low-cost anomaly detection method and system based on multi-channel cameras, which can efficiently and accurately check abnormal situations in real time.
[0012] On the one hand, to achieve the above-mentioned purpose, the present invention provides a low-cost anomaly detection method based on multiple cameras, comprising:
[0013] The multi-channel camera acquisition module collects the field of view images in front of the locations of the multi-channel cameras in real time;
[0014] Perform real-time anomaly detection on the collected images through a lightweight anomaly target detection model and output the detection results;
[0015] Based on the detection results, combined with multi-camera hierarchical detection and multi-field dual-frame detection strategies, the detection mode is dynamically switched, and a hierarchical response of low-risk mode, potential abnormal mode and high abnormal warning mode is achieved by calculating historical anomaly scores.
[0016] Preferably, the layout of the multi-channel cameras includes a composite layout of at least one wide-angle camera and multiple narrow-angle cameras, achieving full coverage of the field of view in front of the cameras through AHD technology, and constructing a multi-channel AHD signal analysis unit using the XS9922B chip, and outputting multiple video streams to the main control platform using virtual channel technology.
[0017] Preferably, the multi-channel cameras are connected to the CSI module of the main control platform through the XS9922B chip to achieve low-latency capture of the original image data, and finally through RTP encapsulation and UDP transmission to achieve remote real-time video transmission.
[0018] Preferably, the lightweight abnormal target detection model is the YOLOv8n-ADL model. The YOLOv8n-ADL model uses GhostConv to replace traditional convolution, generates feature maps through identity mapping and cheap operations; and introduces the LSKA attention mechanism in the SPPF module to capture multi-scale features through separable one-dimensional convolutional kernels; uses the WaveletUnPool module to replace traditional upsampling and restores image details using wavelet filters; uses the LSDECD detection head to share convolution parameters, and uses the MPDIoU loss function to optimize the target box overlap problem.
[0019] Preferably, the multi-camera hierarchical detection includes:
[0020] Based on the detection results of the wide-angle camera, determine whether to enable the narrow-angle camera;
[0021] Perform hierarchical detection on abnormal targets with a confidence level of the abnormal target exceeding a preset threshold, and fuse the confidence scores of the wide-angle camera and the narrow-angle camera to calculate the abnormal situation score. Specifically:
[0022]
[0023] In the formula, score is the abnormal situation score, N c is the number of narrow-angle cameras, C n,i is the set of confidence levels of high-confidence abnormal targets detected by the i-th narrow-angle camera, c i is the target confidence level in C n,i , C w is the set of confidence levels of the wide-angle camera, and c is the target confidence level in C w .
[0024] Preferably, the multi-field-of-view dual-frame detection strategy is:
[0025] Stitch the wide-angle camera image and the narrow-angle camera image selected by polling into a dual-frame input image, dynamically select the detection object according to the time difference detected by the narrow-angle camera last time, traverse the detected abnormal targets and calculate the abnormal score;
[0026] Among them, dynamically selecting the detection object is specifically:
[0027] i = argmax({T i | i ∈ {1, …, N c}});
[0028] In the formula, i is the narrow-angle camera number, T i is the time difference between the narrow-angle camera i and the last call, and N c is the number of narrow-angle cameras.
[0029] Preferably, before performing the multi-camera hierarchical detection and multi-field-of-view dual-frame detection strategy, it also includes: obtaining the mapping relationship between the wide-angle camera and the narrow-angle camera through a calibration module, where the calibration module maps the field of view of the narrow-angle camera to the unified coordinate system of the wide-angle camera through checkerboard calibration and perspective transformation matrix calculation, and determines the projection area of the narrow-angle camera through mask matching.
[0030] Preferably, calculating the historical abnormal score is:
[0031] F(x) = (1 - α) * F(x - 1) + α * score;
[0032] In the formula, α is the weighting coefficient, score is the abnormal score, F(x) is the historical abnormal score at the current moment x, and F(x - 1) is the historical score at the previous moment x - 1.
[0033] On the other hand, to achieve the above object, the present invention also provides a low-cost abnormal detection system based on a multi-channel camera, including:
[0034] Multi-channel camera acquisition module: used to collect the field-of-view images in front of the positions of the multi-channel cameras in real time;
[0035] Lightweight abnormal detection module: used to perform real-time abnormal detection on the collected images through an improved YOLOv8n model and output the detection results;
[0036] Mode switching and abnormal processing module: used to dynamically switch the detection mode according to the detection results in combination with the multi-camera hierarchical detection and multi-field-of-view dual-frame detection strategy, and achieve hierarchical response of low-risk mode, potential abnormal mode and high-abnormal warning mode by calculating the historical abnormal score.
[0037] Preferably, the system supports access of up to 16 AHD cameras, and through the nvv4l2h264enc hardware accelerator, the single-channel encoding delay is ≤9.5 ms and the eight-channel concurrent delay is ≤45.4 ms.
[0038] Compared with the prior art, the present invention has the following advantages and technical effects:
[0039] 1. Cost optimization of multiple cameras in low-speed scenarios: The present invention adopts a multiple-channel low-cost AHD camera access solution. By collocating wide-angle cameras and narrow-distance cameras, it realizes the acquisition of all-round views at different distances. Without sacrificing the performance of the system, it reduces the hardware cost and achieves the effect of cost reduction and efficiency improvement;
[0040] 2. For common pedestrian and environmental abnormal scenarios, improve and lightweight YOLOv8n. Without occupying too much computing resources, it can efficiently and accurately check abnormal situations in real time;
[0041] 3. The present invention proposes a mode switching and anomaly handling module based on historical scores. Through the historical score situation, this mechanism can intelligently switch between multi-view dual-frame detection and multi-camera hierarchical detection, efficiently detect and save resources during long-term low risks; comprehensively screen to avoid missed detections in potential anomalies, and give timely warnings in high anomalies; achieve reasonable resource allocation and efficient and accurate handling of anomalies, effectively utilize computing power, reduce ineffective detection redundancy, quickly respond to and make decisions on abnormal events, and reduce the losses and impacts brought by anomalies. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0043] Figure 1 It is a flowchart of a low-cost anomaly detection method based on multiple cameras according to an embodiment of the present invention;
[0044] Figure 2 It is a schematic diagram of the layout of multiple cameras according to an embodiment of the present invention;
[0045] Figure 3 It is a schematic diagram of the parameters of multiple cameras according to an embodiment of the present invention;
[0046] Figure 4 It is a schematic diagram of the construction process of a lightweight anomaly target detection model according to an embodiment of the present invention;
[0047] Figure 5 It is a flowchart of multi-camera hierarchical detection according to an embodiment of the present invention;
[0048] Figure 6 Schematic diagram of the multi - field - of - view dual - frame detection strategy for the embodiments of the present invention;
[0049] Figure 7 Flowchart of the calibration module for the embodiments of the present invention;
[0050] Figure 8 Flowchart of mode switching and anomaly handling based on historical scores for the embodiments of the present invention. Detailed implementation manners
[0051] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will detail this application with reference to the drawings and in combination with the embodiments.
[0052] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer - executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that here.
[0053] The present invention proposes a low - cost anomaly detection method based on multiple cameras, as Figure 1 , including:
[0054] Collect the field - of - view images in front of the positions where multiple cameras are located in real - time through a multi - camera acquisition module;
[0055] Perform real - time anomaly detection on the collected images through an improved YOLOv8n model and output the detection results;
[0056] Based on the detection results, combine the multi - camera hierarchical detection and multi - field - of - view dual - frame detection strategies, dynamically switch the detection mode, and achieve hierarchical responses in low - risk mode, potential anomaly mode, and high - anomaly warning mode by calculating the historical anomaly scores.
[0057] In this embodiment, for low - speed scenarios, combined with edge computing nodes such as Jetson AGX Orin, a low - cost multi - camera device is designed to obtain an all - around field of view at different distances, achieve a full coverage of the 180° front - view field, and reduce the hardware cost of the multi - camera system; for the processing of abnormal scenarios, a lightweight model is combined to quickly and accurately detect the inputs of multiple cameras; at the same time, a mode - switching and anomaly - handling method based on historical scores is adopted, an efficient detection strategy suitable for running in abnormal and non - abnormal modes is proposed, and mode switching and anomaly handling are performed according to the historical anomaly scores of the current system. While comprehensively detecting each camera, the number of ineffective detections is greatly reduced.
[0058] Furthermore, the layout of the multi-channel cameras includes a composite layout of at least one wide-angle camera and multiple narrow-angle cameras. The AHD technology is used to achieve full coverage of the field of view in front of the cameras' positions. An XS9922B chip is used to construct a multi-channel AHD signal parsing unit, and the virtual channel technology is utilized to multiplex multiple video streams into a single MIPI-CSI2 interface and output it to the Jetson AGX Orin main control platform.
[0059] The multi-channel cameras are connected to the CSI module of the main control platform through the XS9922B chip to achieve low-latency capture of the original image data. The nvv4l2h264enc hardware accelerator is adopted to achieve encoding with 9.5 ms for a single 1080P@25fps video and 45.4 ms for eight-channel concurrency. Finally, through RTP encapsulation and UDP transmission, remote real-time video transmission is realized.
[0060] Specifically, as Figures 2 - 3 , in this embodiment, the multi-channel cameras use IMX290 as the image sensor of the multi-channel AHD cameras. Its AHD signal output ability of 1080P@25fps, combined with the characteristics of wide dynamic range and high sensitivity, can adapt to complex lighting environments. The system realizes full coverage of the 180° far and near detection distance of the front field of view through the composite layout of a 180° wide-angle lens and multiple narrow-angle cameras.
[0061] An XS9922B chip is used to construct a 4-channel AHD signal parsing unit. A single XS9922B chip supports the parsing of 4-channel 1080P@30fps AHD video signals (compatible with HDCCTV and CVBS protocols). Through the virtual channel (VC) technology, signal-level multiplexing of four independent video streams is realized, and they are aggregated into a single MIPI-CSI2 interface output through a four-channel MIPI TX interface (each channel supports a maximum transmission rate of D-PHY 1.5Gbps). The main control platform Jetson AGX Orin realizes a maximum video input capacity of 16 channels through a 4*4lane MIPI-CSI2 interface.
[0062] A high-performance video processing pipeline is built based on the Jetson AGX Orin platform. By connecting the CSI module to the MIPI-CSI2 channel of the XS9922B, low-latency capture of the original image data is achieved; the VI module further processes the original data and supports the parsing and transmission of up to 16 channels of video through the virtual channel management mechanism; the encoding link adopts the nvv4l2h264enc hardware accelerator to achieve fast encoding with 9.5 ms for a single 1080P@25fps video and 45.4 ms for eight-channel concurrency; finally, through RTP encapsulation and UDP transmission, remote real-time video transmission is realized, supporting real-time video acquisition and analysis at the remote end.
[0063] Through the AHD solution, in low-speed scenarios, efficient access to up to 16 cameras is achieved, which greatly reduces the cost of the multi-camera system while ensuring the scene requirements. By selecting a composite layout of at least one 180° wide-angle lens and multiple narrow-angle cameras, it is ensured that the narrow-angle cameras can jointly cover a 180° field of view in the front, and their physical positions are fixed to ensure that the relative physical positions of the narrow-angle cameras and the wide-angle camera remain unchanged, thereby achieving full coverage of the detection distance of the 180° far and near front view. In this embodiment, the number of narrow-angle cameras is three.
[0064] Furthermore, the lightweight anomaly target detection model is the YOLOv8n-ADL model. The YOLOv8n-ADL model uses GhostConv to replace the traditional convolution, generates feature maps through identity mapping and cheap operations; and introduces the LSKA attention mechanism in the SPPF module to capture multi-scale features through separable one-dimensional convolutional kernels; uses the WaveletUnPool module to replace the traditional upsampling and restores image details with wavelet filters; uses the LSDECD detection head to share convolution parameters, and uses the MPDIoU loss function to optimize the target box overlap problem.
[0065] Specifically, as Figure 4 , based on the improvement of the YOLOv8n model, this embodiment proposes a lightweight anomaly target detection model, YOLOv8n-ADL (YOLOv8n - Anomaly Detect - Lightweight), which includes steps of data collection and processing, model improvement and lightweighting, model training, and model performance evaluation.
[0066] Step 1: Data collection and processing are divided into data acquisition, data annotation, and dataset division:
[0067] Data collection: Based on common requirements in public safety, this embodiment constructs an anomaly event dataset covering environmental anomalies and pedestrian behavior anomalies. Among them, the environmental anomaly categories are: fire, smoke, and traffic accidents; pedestrian behavior anomalies are mainly falls, fights, and carrying weapons. Since there are few anomaly situation datasets in reality, when making the dataset, the datasets on the network are crawled; especially for the pedestrian anomaly dataset, the abnormal behavior videos collected on the network are used, and images are collected every 5 frames to effectively solve the oversampling problem caused by action redundancy. Through the standardized sampling process, 2500 high-quality samples are obtained for each type of anomaly scene, and finally an anomaly event dataset of 15000 RGB images is constructed.
[0068] Data annotation: Using the LabelImg software, 15,000 datasets are annotated, and the image annotations will be saved in the TXT file format. Six types of targets to be saved are annotated, namely fire, smoke, traffic accident, fall, fight, and weapon-hold.
[0069] Dataset division: The stratified random sampling strategy is adopted for dataset division. According to the ratio of 4:1, the dataset is divided into a training set and a validation set. Finally, 12,000 images are obtained for the training set and 3,000 images for the validation set.
[0070] Step 2: Model improvement and lightweighting:
[0071] Aiming at the problems of insufficient local feature capture, decreased sensitivity to multi-scale targets of the YOLOv8n baseline model in complex scenarios, and in order to further reduce the consumption of system resources, this embodiment improves and lightweights the YOLOv8n network.
[0072] First, GhostConv is used to replace the traditional convolution, and C3Ghost is used to replace the traditional C2f module, reducing the number of model parameters without sacrificing detection accuracy. Secondly, the LSKA attention mechanism is introduced into SPFF to efficiently obtain multi-scale features. Then, the WaveletUnPool module is used to replace the traditional upsampling module, which is beneficial for the model to restore image details in the upsampling process and improve the recognition accuracy of small targets. Then, the LSDECD head is used to replace the detection head, reducing the number of model parameters while sharing convolution parameters. Finally, the MPDIoU loss function is used to improve the problem of the overlap of human target boxes.
[0073] The specific modules are described as follows:
[0074] Module 1, GhostConv module: At the 1st, 3rd, 5th, and 7th layers of the Backbone, GhostConv is adopted. GhostConv is specifically divided into two steps. First, standard convolution is used to generate the initial feature map:
[0075] Y = X × f + b;
[0076] where X is the input feature, f is the convolution kernel, b is the bias constant, and Y is the output feature map.
[0077] Then, for each channel's feature map y' of the output feature map Y i , φ i,j is used to generate the Ghost feature map y' i,j , and finally the Ghost feature map is obtained:
[0078]
[0079] Among them, φ i,j contains an identity mapping. After performing s mappings on it, s mappings include one identity mapping, and s - 1 is the cheap operate operation. Finally, a channel output result of m * s is obtained.
[0080] Module 2, C3Ghost module: At the 2nd, 4th, 6th, and 8th layers of the Backbone, the C3Ghost module is adopted. This module contains the GhostBottleneck module. The GhostBottleneck module sequentially performs GhostConv, batch normalization (BN), SiLU activation function, and GhostConv, and then performs an addition (Add) operation on the result and the original feature map to reduce the number of parameters while maintaining the representational ability.
[0081] Module 3, SPPF_LSKA module: The SPPF module has room for improvement in capturing small and fine-grained information and is difficult to adapt to complex abnormal scene detection. Therefore, at the 9th layer of the Backbone, the LSKA attention module is introduced and combined with the original SPPF module to effectively improve the ability of the backbone feature network to extract features at multiple scales, thereby improving the detection effect. The SPPF_LSKA module decomposes the traditional two-dimensional weight convolution kernel into two cascaded one-dimensional separable convolution kernels, greatly reducing the computational complexity and memory requirements, and can capture a wider range of image features while maintaining efficient computing.
[0082] Module 4, WaveletUnPool module: At the 10th and 13th layers of the Neck, the WaveletUnPool upsampling module is adopted. Four wavelet filters of size 2×2 are defined, namely low-frequency - low-frequency, low-frequency - high-frequency, high-frequency - low-frequency, and high-frequency - high-frequency, which are used to process the smooth part, vertical direction, horizontal direction, and image detail part of the image. These wavelet filters are used to better restore the image details through the deconvolution (i.e., unpooling) operation, which helps the model capture the details of small targets more accurately.
[0083] Module 5, LSDECD: In the Head, LSDECD is used to replace the original head of YOLO8n. LSDECD combines the original three feature extractions in the head into a shared convolution, then uses a scale scaling module for scale scaling, and replaces all the normalization layers in the convolution with Group Normalization (GN) to ensure that the model has sufficient localization and classification performance while reducing the model parameters.
[0084] Module 6, the MPDIoU loss function, can be used to reduce the problem of decreased recognition performance caused by overlapping pedestrian bounding boxes.
[0085] The MPDIoU calculation formula is as follows:
[0086]
[0087] Among them, d1 and d2 are the squares of the Euclidean distances from the upper left corner point of the predicted bounding box to the upper left corner point and the lower right corner point of the ground truth bounding box respectively; w and h represent the width and height of the ground truth bounding box respectively, IoU is the intersection over union of the predicted box and the ground truth box, and MPDIoU can reduce the problem of decreased recognition performance caused by overlapping pedestrian bounding boxes.
[0088] Step 3: Model training:
[0089] This object recognition algorithm is under the Ubuntu20.04 LTS operating system, using the PyTorch2.0.0+cu124 deep learning framework, the python3.10.12 environment, with the hardware configuration of an Intel(R) Core(TM) i5-13600KF@3.5GHz (16-core, 24-thread) processor and an NVIDIA GeForce RTX 4060Ti 8GB CDDR6 graphics card. The YOLOV8n-ADL model is trained, the number of training epochs is 300, the SGD optimizer is used for gradient update, the batch-size is 32, the initial learning rate lr0 is 0.001, and the trained weight file best.pt is saved.
[0090] Model performance evaluation: To objectively evaluate the performance of the object detection algorithm, precision P, recall R, average precision AP, number of parameters Param, floating point operations per second FLOPs, and model memory occupancy are used as evaluation metrics. As shown in Table 1.
[0091] Table 1
[0092]
[0093] Furthermore, multi-camera hierarchical detection includes:
[0094] Based on the detection results of the wide-angle camera, determine whether to enable the narrow-angle camera;
[0095] Perform hierarchical detection on abnormal objects whose confidence levels exceed a preset threshold, and fuse the confidence scores of the wide-angle camera and the narrow-angle camera to calculate the abnormal situation score. Specifically:
[0096]
[0097] Where score is the anomaly score, and N c is the number of narrow-angle cameras, and C n,i is the set of confidences of high-confidence anomaly targets detected by the i-th narrow-angle camera, and c i is the target confidence in C n,i , and C w is the set of confidences of the wide-angle camera, and c is the target confidence in C w .
[0098] Specifically, multi-camera hierarchical detection means that within the same acquisition cycle, the system will synchronously acquire a frame of wide-angle camera image and images of each narrow-angle camera. Since there are mapping regions of each narrow-angle camera in the wide-angle camera, a hierarchical detection strategy can be used to mine the relevant information between the wide-angle camera and the images of each narrow-angle camera to accurately detect abnormal situations.
[0099] For example Figure 5 , the specific steps of multi-camera hierarchical detection are as follows:
[0100] Step 1: Detect the wide-angle camera and determine whether to enable the narrow-angle camera for detection.
[0101] Using the YOLOv8n-ADL algorithm, first detect abnormal targets in the wide-angle camera. If a target is detected, traverse the N detected targets to obtain the confidence Conf j of each target and the target box Box j (1 ≤ j ≤ N).
[0102] For each target j, calculate whether it is within the mapping field of view of the narrow-angle camera i (1 ≤ j ≤ N c ), where N c is the number of narrow-angle cameras, and the calculation formula is as follows:
[0103]
[0104] where R i represents the mapping region of the narrow-angle camera i in the wide-angle camera, InRegion j is the mapping region number where the j-th target is located, and N c is the number of narrow-angle cameras;
[0105] If InRegion j ≠ 0, it indicates that the j-th target is located in the narrow-angle camera i, and the box area of the j-th target does not exceed the imaging field of view of the narrow-angle camera i. At this time, according to the result of InRegion j and the confidence of the j-th target, determine whether to turn on the narrow-angle camera i for detection:
[0106]
[0107] Among them, Conf low is the set low confidence threshold, is the flag indicating whether the narrow-angle camera i (i = InRegion j ) needs further detection;
[0108] When the target j is within the imaging field of view of the narrow-angle camera i and the confidence level exceeds Conf low , set to 1, and it is necessary to further detect the narrow-angle camera i to prevent missed detection of abnormal conditions.
[0109] Step 2: Save the abnormal target threshold of the wide-angle camera:
[0110] Traverse the N detected abnormal targets, and classify the targets into two categories for saving according to whether the target is within the imaging field of view of the narrow-angle camera i and the different confidence levels:
[0111] Set C1 w : The target is not within the imaging field of view of the narrow-angle camera i and the confidence level exceeds Conf low :
[0112] C1 w ={Conf j |j ∈ [1, N], InRegion j ≠ 0 and Conf j > Conf low};
[0113] Set C2 w : The target is within the imaging field of view of the narrow-angle camera i and the confidence level exceeds Conf high :
[0114] C2 w ={Conf j |j ∈ [1, N], InRegion j == 0 and Conf j > Conf high};
[0115] Merge set C1 w and C2 w , to obtain the confidence level set C of all abnormal targets for the wide-angle camera w :
[0116] C w = C1 w ∪ C2 w .
[0117] Step 3: Enable the narrow-angle camera i for detection and save the results:
[0118] Traverse the abnormal flag Abnormal corresponding to each narrow-angle camera i i , if Abnormal i == 1, then call YOLOv8n-ADL to detect the narrow-angle camera i.
[0119] Traverse the N i detected abnormal targets, record the confidence as and save the targets with confidence exceeding Conf low :
[0120]
[0121] Step 4: Calculate the abnormal score:
[0122] The abnormal score is obtained by accumulating the confidence values of all elements in the confidence set C w of the wide-angle camera and the confidence values of all elements in the confidence set C n,i of all narrow-angle cameras. The calculation formula is as follows:
[0123]
[0124] In the formula, score is the abnormal situation score, N c is the number of narrow-angle cameras, C n,i is the confidence set of high-confidence abnormal targets detected by the i-th narrow-angle camera, c i is the target confidence in C n,i , C w is the confidence set of the wide-angle camera, and c is the target confidence in C w .
[0125] Furthermore, the multi-field double-frame detection strategy is:
[0126] Stitch the wide-angle camera image and the narrow-angle camera image selected by polling into a double-frame input image, dynamically select the detection object according to the time difference of the previous detection of the narrow-angle camera, traverse the detected abnormal targets and calculate the abnormal score;
[0127] Among them, dynamically selecting the detection object is specifically:
[0128] i = argmax({T i | i ∈ {1, …, N c}});
[0129] In the formula, i is the narrow-angle camera number, Ti The time difference between the last call of the narrow-angle camera i and N c is the number of narrow-angle cameras.
[0130] Specifically, the multi-field-of-view dual-frame detection strategy is designed on the basis of fully considering the input requirements of the YOLO model and the characteristics of the AHD image resolution, aiming to optimize the utilization efficiency of image detection resources, ensure the detection effect, and achieve comprehensive and efficient detection of multi-camera data. Through its unique dual-frame detection strategy and multi-field-of-view fusion mechanism, while reducing image resource waste, it can take into account a large range of fields of view and distant detailed information for comprehensive detection.
[0131] The following is a further detailed elaboration of its working principle:
[0132] Image resource optimization:
[0133] Since the YOLO model requires a square image as input for detection, conventional AHD images (common resolutions are (1920, 1080), (1280, 720)) need to be padded before being input into the model to adjust to a square image with an aspect ratio of 1:1, which will cause a large amount of image resource waste. For an AHD image with a resolution of (w, h), after padding, it will waste the image resources. The dual-frame detection strategy forms an image of size w×(2×h) by splicing two frames of images vertically and then sends it into the model. In this way, two frames of images can be detected simultaneously in one detection, and for each frame of image, only the image space will be wasted. This strategy can achieve the detection of multi-cameras with less system resource consumption. the image space, and this strategy can detect multiple cameras with less system resource consumption.
[0134] Multi-field-of-view collaborative detection:
[0135] The wide-angle camera can capture a large range of fields of view, which is of great significance for the overall scene perception. The module will continuously detect the wide-angle camera to obtain the overall situation of the entire monitoring area in real time and promptly discover possible abnormal events; on the basis of ensuring continuous detection of the wide-angle camera, it will detect each narrow-angle camera in a polling manner according to a certain strategy. Narrow-angle cameras usually focus on specific areas or objects and can provide more detailed information. Through the polling detection method, it can ensure that abnormal events in the distance will not be missed.
[0136] Such as Figure 6 The specific steps are as follows:
[0137] Step 1: In the same acquisition cycle, the system will synchronously acquire one frame of the wide-angle camera's image and the images of all narrow-angle cameras.
[0138] On the one hand, the dual-frame detection strategy will fixedly select the image of the wide-angle camera as the first frame. On the other hand, it will determine based on the time difference T between each narrow-angle camera and the last detection. i The selection formula is as follows:
[0139] i = arg max({T i | i ∈ {1, …, N c}});
[0140] Through the above formula, the narrow-angle cameras will be selected regularly and in a polling manner according to the time sequence of the last detection on each narrow-angle camera, which can ensure that each narrow-angle camera can be detected fairly.
[0141] Step 2: Invoke the YOLOv8n-ADL algorithm to detect the image after combining the dual frames. Traverse the N d detected abnormal targets, record the confidence level as ConfDouble jd (1 ≤ jd ≤ N d ) and save the targets with a confidence level exceeding Conf low :
[0142] C d = {ConfDouble jd | jd ∈ [1, N d , ConfDouble jd > Conf low};
[0143] Step 3: Calculate the anomaly score:
[0144] The anomaly score is accumulated from the elements of the confidence set C d of the image after combining the dual frames:
[0145]
[0146] In the formula, score is the anomaly score, and cd is the target confidence level in C d .
[0147] Furthermore, before performing multi-camera hierarchical detection and multi-field-of-view dual-frame detection strategy, it also includes: obtaining the mapping relationship between the wide-angle camera and the narrow-angle camera through a calibration module. Among them, the calibration module maps the field of view of the narrow-angle camera to the unified coordinate system of the wide-angle camera through checkerboard calibration and perspective transformation matrix calculation, and determines the projection area of the narrow-angle camera through mask matching.
[0148] Specifically, the calibration module functions in the system initialization stage. Since the relative physical positions of the multi-channel narrow-angle cameras and the wide-angle camera remain unchanged, it is only necessary to obtain the mapping relationship between the wide-angle camera and each narrow-angle camera once during system initialization.
[0149] Such as Figure 7 , specifically including:
[0150] Step 1: Through checkerboard calibration, obtain the distortion coefficients of the wide-angle camera and perform distortion correction on the wide-angle camera. Denote the field of view obtained after the distortion correction of the wide-angle camera as S0.
[0151] Step 2: Jointly calibrate each narrow-angle camera and the wide-angle camera. Use the calibration template to mark the field of view of the narrow-angle camera i as Si, mark the four vertices of the narrow-angle camera template and the corresponding four vertices in the wide-angle camera. Use these corresponding vertices to obtain the perspective transformation matrix H from the narrow-angle camera to the wide-angle camera:
[0152]
[0153] Among them, represents a linear transformation, such as scaling, shearing, and rotation. [a 31 a 32 is used for translation. [a 13 a 23 T is used to generate a perspective transformation. Use this transformation matrix to map the image of the narrow-angle camera to the same perspective as the wide-angle camera:
[0154] [x ′ , y ′ , z′] = H[x, y, z];
[0155] Among them, x, y, z are the original coordinate values of the narrow-angle camera, and x ′ , y ′ , z′ are the coordinate values of the narrow-angle camera after perspective transformation.
[0156] Step 3: Calculate the projection area based on mask-based template matching.
[0157] Since the area after perspective transformation is usually not rectangular, a mask is used to mark the background of the image of the narrow-angle camera after perspective transformation. Mask-based template matching is used to find the projection area of the narrow-angle camera in the wide-angle camera and record it as Ri.
[0158] Furthermore, calculate the historical anomaly score as:
[0159] F(x) = (1 - α) * F(x - 1) + α * score;
[0160] Where α is the weighting coefficient, score is the anomaly score, F(x) is the historical anomaly score at the current moment x, and F(x - 1) is the historical score at the previous moment x - 1.
[0161] Specifically, as Figure 8 , by defining the historical anomaly score formula, the dynamic switching of the detection mode and anomaly handling are realized according to the anomaly historical data. The specific process is as follows:
[0162] Step 1: Use the calibration module to calibrate the wide-angle camera and each narrow-angle camera to obtain the projection area of each narrow-angle camera in the imaging image of the wide-angle camera;
[0163] Step 2: To reasonably measure the historical situation of the system's abnormal state, define the historical abnormal situation score formula as follows:
[0164] F(x) = (1 - α) * F(x - 1) + α * score;
[0165] Among them, α is the weighting coefficient, which is used to adjust the balance between the historical score and the current score. Usually, to ensure more emphasis on the current detection situation, α is usually taken as α > 0.5; score is the abnormal score situation of this round output by both the multi-view double-frame detection mode and the multi-camera hierarchical detection mode during operation. F(x) is the historical anomaly score at the current moment x, and F(x - 1) is the historical score at the previous moment x - 1.
[0166] Step 3: Obtain images of the wide-angle camera and each camera at the same moment;
[0167] Step 4: Mode switching and processing strategy based on the historical abnormal situation:
[0168] 1. Long-term low-risk mode switching: When F(x) < F low (F low ≈ 0), F low is the low-risk threshold of the abnormal score, indicating that the system has been in a stable non-abnormal situation for a long time. In this case, even if misdetection occurs occasionally, the system can quickly adjust according to its historical score. At this time, the multi-view double-frame detection mode is selected. This mode can efficiently detect each camera, make full use of the high efficiency and low resource consumption of the double-frame detection strategy, and realize the efficient use of system resources while ensuring detection.
[0169] 2. Potential anomaly comprehensive detection mode switching: When F(x) > F lowWhen this happens, it means that some abnormal situations may have occurred in the system. To comprehensively and accurately detect potential abnormal conditions and prevent missed detections, the module will switch to a multi-camera hierarchical detection mode at this time. This mode conducts preliminary screening starting from the wide-angle camera through hierarchical detection, and then specifically enables the narrow-angle camera for fine detection according to the situation, so as to achieve comprehensive coverage and in-depth detection of the entire monitoring scene.
[0170] 3. High Abnormality Alarm and Early Warning Mode Processing: When F(x) > F high When high F is the high-risk threshold of the abnormality score, and the system is clearly in a significant abnormal situation. In this case, not only will it trigger the local device to issue an abnormality alarm, but it will also promptly send early warning information to the remote server and synchronously transmit the real-time video stream. This helps relevant personnel receive the abnormal information in a timely manner, make quick responses and decisions, and minimize the losses and impacts that may be brought by abnormal events.
[0171] Step 5: Return to Step 3.
[0172] This embodiment also provides a low-cost abnormal detection system based on multiple cameras, including:
[0173] Multiple Camera Acquisition Module: Used to collect the field-of-view images in front of the positions of multiple cameras in real time;
[0174] Lightweight Abnormal Detection Module: Used to perform real-time abnormal detection on the collected images through an improved YOLOv8n model and output the detection results;
[0175] Mode Switching and Abnormality Processing Module: Used to dynamically switch the detection mode according to the detection results in combination with the multi-camera hierarchical detection and multi-field-of-view dual-frame detection strategies, and achieve hierarchical responses in the low-risk mode, potential abnormal mode, and high-abnormality early warning mode by calculating the historical abnormality scores.
[0176] The system supports up to 16 AHD cameras to be connected, and through the nvv4l2h264enc hardware accelerator, the single-channel encoding delay ≤ 9.5ms and the eight-channel concurrent delay ≤ 45.4ms are achieved.
[0177] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A low-cost anomaly detection method based on multiple cameras, characterized in that, Including: The multi-channel camera acquisition module is used to collect the vision images in front of the positions of multiple cameras in real time; The lightweight abnormal target detection model is used to perform real-time abnormal detection on the collected images and output the detection results; Based on the detection results, combined with the multi-camera hierarchical detection and multi-field double-frame detection strategies, the detection mode is dynamically switched, and the hierarchical response of the low-risk mode, potential abnormal mode and high-abnormal warning mode is realized by calculating the historical abnormal score.
2. The low-cost anomaly detection method based on multiple cameras according to claim 1, wherein The layout of the multi-channel cameras includes a composite layout of at least one wide-angle camera and multiple narrow-distance cameras. The AHD technology is used to achieve full coverage of the vision in front of the positions of the cameras. An XS9922B chip is used to construct a multi-channel AHD signal analysis unit, and the virtual channel technology is used to output multiple video streams to the main control platform.
3. The low-cost anomaly detection method based on multiple cameras according to claim 2, characterized in that, The multi-channel cameras are connected to the CSI module of the main control platform through the XS9922B chip to achieve low-latency capture of the original image data. Finally, through RTP encapsulation and UDP transmission, remote real-time video transmission is realized.
4. The low-cost anomaly detection method based on multiple cameras according to claim 1, characterized in that, The lightweight abnormal target detection model is the YOLOv8n-ADL model. The YOLOv8n-ADL model uses GhostConv to replace the traditional convolution, and generates feature maps through identity mapping and cheap operations; and the LSKA attention mechanism is introduced into the SPPF module to capture multi-scale features through separable one-dimensional convolution kernels; the WaveletUnPool module is used to replace the traditional upsampling, and the wavelet filter is used to restore the image details; the LSDECD detection head shares convolution parameters, and the MPDIoU loss function is used to optimize the target box overlap problem.
5. The low-cost anomaly detection method based on multi-channel cameras according to claim 1, characterized in that The multi-camera hierarchical detection includes: Based on the detection results of the wide-angle camera, judge whether to enable the narrow-distance camera; Perform hierarchical detection on abnormal targets whose confidence levels of abnormal targets exceed the preset threshold, and fuse the confidence scores of the wide-angle camera and the narrow-distance camera to calculate the abnormal situation score. Specifically: Where score is the anomaly score, N c is the number of narrow-angle cameras, C n,i is the set of confidences of high-confidence anomaly targets detected by the i-th narrow-angle camera, c i is the target confidence in C n,i , C w is the set of confidences of the wide-angle camera, and c is the target confidence in C w .
6. The low-cost anomaly detection method based on a multi-channel camera according to claim 1, characterized in that The multi-field double-frame detection strategy is: Stitch the wide-angle camera image and the narrow-distance camera image selected by polling into a double-frame input image, and dynamically select the detection object according to the time difference of the last detection of the narrow-distance camera, traverse the detected abnormal targets and calculate the abnormal score; Among them, dynamically selecting the detection object is specifically: i = argmax({T i | i ∈ {1, …, N c}}); Where i is the number of narrow - distance cameras, and T i is the time difference between the last call of narrow - distance camera i, and N c is the number of narrow - distance cameras.
7. The low-cost anomaly detection method based on multiple cameras according to claim 1, characterized in that, Before performing the multi-camera hierarchical detection and multi-field double-frame detection strategies, it also includes: obtaining the mapping relationship between the wide-angle camera and the narrow-distance camera through the calibration module. Among them, the calibration module calculates through checkerboard calibration and perspective transformation matrix, maps the field of view of the narrow-distance camera to the unified coordinate system of the wide-angle camera, and determines the projection area of the narrow-distance camera through mask matching.
8. The low-cost anomaly detection method based on multiple cameras according to claim 1, wherein Calculating the historical abnormal score is: F(x) = (1 - α) * F(x - 1) + α * score; In the formula, α is the weighting coefficient, score is the abnormal score, F(x) is the historical abnormal score at the current moment x, and F(x - 1) is the historical score at the previous moment x - 1.
9. A low-cost anomaly detection system based on multiple cameras, characterized in that, Including: Multi-channel camera acquisition module: used to collect the vision images in front of the positions of multiple cameras in real time; Lightweight anomaly detection module: It is used to perform real-time anomaly detection on the collected images through an improved YOLOv8n model and output the detection results; Mode switching and anomaly handling module: It is used to dynamically switch the detection mode according to the detection results in combination with the multi-camera hierarchical detection and multi-field-of-view dual-frame detection strategies, and achieve hierarchical responses in low-risk mode, potential anomaly mode and high-anomaly warning mode by calculating the historical anomaly scores.
10. The low-cost anomaly detection system based on multiple cameras according to claim 9, wherein The system supports access of up to 16 AHD cameras, and realizes a single-channel encoding delay ≤ 9.5 ms and an eight-channel concurrent delay ≤ 45.4 ms through the nvv4l2h264enc hardware accelerator.
Citation Information
Patent Citations
Multi-modal target detection method used in complex scene
CN116630608A
Image processing method and system for intelligent security and protection monitoring
CN118887622A
Video monitoring safety control system and method based on Internet of Things
CN119254924A
Line cutter
KR1020210033162A
Cited By
Multi-camera network optimization selection method based on large model
CN121174052A
Lightweight multi-scale SAR image target detection method based on improved YOLO architecture
CN121746691A