Monitoring video detection method and device based on array camera, equipment and medium
By using panoramic and detailed image stitching technology from array cameras, the problem of low accuracy in anomaly detection in surveillance videos has been solved, enabling efficient and automatic anomaly detection in surveillance videos and adapting to the needs of different monitoring levels.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN SUPERNODE NETWORK TECH
- Filing Date
- 2025-09-09
- Publication Date
- 2026-04-21
AI Technical Summary
Existing surveillance video has low accuracy in detecting anomalies, manual monitoring is costly and prone to fatigue, it is difficult to integrate third-party software or hardware into existing surveillance systems for automatic monitoring, and ordinary cameras have low resolution and require a large number of installations.
An array camera, including a wide-angle lens and multiple detail lenses, is used to calculate the homography matrix by feature detection and stitching of panoramic and regional detail images, thereby achieving high-resolution detection of surveillance videos.
It improves the accuracy of anomaly detection in surveillance videos, reduces monitoring costs, and enables automatic monitoring without the need for software installation, adapting to the needs of different monitoring levels.
Smart Images

Figure CN121907987A_ABST
Abstract
Description
[0001] The application is a divisional application with application number 2025112803894 and application date of September 9, 2025, entitled "Method, Apparatus, Device and Medium for Detecting Surveillance Video Based on Array Camera". Technical Field
[0002] This invention relates to the field of image processing, and more particularly to a method, apparatus, computer device, and computer-readable storage medium for detecting surveillance video based on an array camera. Background Technology
[0003] With increasing security demands, video surveillance systems have become a crucial means of maintaining social, business, and public safety. Currently, this typically involves manual monitoring of multiple feeds simultaneously. However, with the proliferation of surveillance cameras, a single person often needs to monitor 10-20 feeds, or even more. This manual monitoring is not only costly but also prone to inattention and fatigue, leading to a decrease in the accuracy of anomaly detection in the surveillance videos. Therefore, improving the accuracy of anomaly detection in surveillance videos has become a pressing technical challenge. Summary of the Invention
[0004] The main objective of this invention is to provide a method, apparatus, device, and computer-readable storage medium for detecting surveillance videos based on an array camera, aiming to improve the accuracy of anomaly detection in surveillance videos.
[0005] To achieve the above objectives, the present invention provides a surveillance video detection method based on an array camera, wherein the array camera includes one wide-angle lens and n detail lenses, and the surveillance video detection method includes the following steps: The number of events n is determined based on the monitoring level of the abnormal events to be detected, and the number of events n is proportional to the monitoring level. The panoramic image is captured through the wide-angle lens, and the detailed image of the region is captured through the n detail lenses, where n is an integer not less than 1; Feature detection is performed on the panoramic image and each regional detail image respectively, and the feature points corresponding to the panoramic image and each regional detail image are calculated as the target matching point set corresponding to the panoramic image and each regional detail image. The homography matrix corresponding to the target matching point set is calculated as the transformation matrix corresponding to each regional detail image on the panoramic image. Based on each transformation matrix, the detailed images of each region are stitched together to obtain a stitched image, which serves as the detailed magnified image corresponding to the panoramic image. The presence of abnormal events in the surveillance video is then detected based on the detailed magnified image.
[0006] Furthermore, to achieve the above objectives, the present invention also provides a surveillance video detection device based on an array camera, wherein the array camera includes one wide-angle lens and n detail lenses, and the surveillance video detection device based on the array camera includes: The detail shot determination module is used to determine n based on the monitoring level of the abnormal event to be detected, and the number of n is proportional to the monitoring level. The terminal screen monitoring module is used to acquire the panoramic image through the wide-angle lens and to acquire the regional detail image through the n detail lenses, where n is an integer not less than 1; The transformation matrix calculation module is used to perform feature detection on the panoramic image and each regional detail image respectively, calculate the feature points corresponding to the panoramic image and each regional detail image as the target matching point set corresponding to the panoramic image and each regional detail image, and calculate the homography matrix corresponding to the target matching point set as the transformation matrix corresponding to each regional detail image on the panoramic image. The abnormal event monitoring module is used to stitch together the detailed images of each region based on various transformation matrices to obtain a stitched image, which serves as a magnified detail image corresponding to the panoramic image, and to detect whether there are abnormal events in the monitoring video based on the magnified detail image.
[0007] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a processor, a memory, and a surveillance video detection program based on an array camera stored in the memory and executable by the processor, wherein when the surveillance video detection program based on an array camera is executed by the processor, the steps of the surveillance video detection method based on an array camera as described above are implemented.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a surveillance video detection program based on an array camera, wherein when the surveillance video detection program based on an array camera is executed by a processor, it implements the steps of the surveillance video detection method based on an array camera as described above.
[0009] This application provides a surveillance video detection method based on an array camera. The array camera includes one wide-angle lens and n detail lenses. The method determines n based on the monitoring level of the abnormal event to be detected, where the number of n is proportional to the monitoring level. A panoramic image is acquired through the wide-angle lens, and regional detail images are acquired through the n detail lenses, where n is an integer not less than 1. Feature detection is performed on both the panoramic image and each regional detail image to calculate the corresponding feature points, which are used as the target matching point set. The homography matrix corresponding to the target matching point set is calculated as the transformation matrix corresponding to each regional detail image on the panoramic image. Based on the transformation matrices, the regional detail images are stitched together to obtain a stitched image, which serves as the magnified detail image corresponding to the panoramic image. The presence of abnormal events in the surveillance video is detected based on the magnified detail image. Through this method, this application acquires surveillance video displayed on a terminal screen using an array camera. By using panoramic and regional detail images of the surveillance video, the system magnifies the details of each frame, improving the image resolution. Anomaly detection is then performed on these high-resolution images, increasing the accuracy of anomaly detection in the surveillance video. Therefore, anomaly detection can be performed on surveillance videos from any terminal without installing any software, significantly enhancing the convenience of surveillance video monitoring. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating the first embodiment of the surveillance video detection method based on an array camera according to the present invention. Figure 2 A schematic diagram of a surveillance video detection device based on an array camera provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the hardware structure of the computer device involved in the embodiment of the present invention.
[0011] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the described order. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0014] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0015] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0016] The surveillance video detection method based on array cameras involved in the embodiments of the present invention is mainly applied to computer equipment, which can be a PC, a portable computer, a mobile terminal or other device with display and processing functions.
[0017] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0018] Reference Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the surveillance video detection method based on an array camera according to the present invention.
[0019] like Figure 1 As shown, this embodiment of the invention provides a surveillance video detection method based on an array camera, which includes steps S10 to S40.
[0020] Step S10: Using an array camera, acquire the monitoring video from at least one terminal screen, and acquire a panoramic image and at least one area detail image of the monitoring video; Currently, in addition to the reduced accuracy of anomaly detection in surveillance videos due to manual monitoring, the following problems also exist: 1. In many high-security monitoring scenarios, it is not permissible to directly integrate any third-party monitoring software or hardware onto existing monitoring systems. In other words, it is not possible to obtain monitoring screen images through software interfaces, software or hardware screenshots, and then allow AI software to perform automatic monitoring.
[0021] 2. If ordinary cameras are used to monitor the surveillance footage, given their relatively low resolution, the cameras need to be installed very close to the monitor screens, and each camera can only monitor one screen. When a monitoring room has dozens or more screens, dozens or more cameras need to be installed accordingly, which is quite difficult to implement.
[0022] To address the aforementioned issues, this application provides a surveillance video detection method based on array cameras. This method enables real-time anomaly monitoring of surveillance videos on terminals where software installation is inconvenient (e.g., in enterprises where complex approval processes hinder software installation or terminal device resources cannot support software installation). By deploying array cameras around the terminal screen, the method captures surveillance videos from the terminal screen, detects abnormal events in the surveillance videos, and promptly identifies and warns of various anomalies based on the detection results, thereby improving monitoring efficiency and security.
[0023] Specifically, the array camera includes at least one wide-angle lens and multiple detail lenses. By deploying multiple high-resolution cameras around each monitoring device in the monitoring room, comprehensive coverage of the monitoring video on all terminal screens is ensured. When deploying the array camera, it is necessary to ensure that the detail images captured by each detail lens are on the panoramic image, and that there are overlapping areas between the detail images captured by each detail lens. Based on the overlapping areas and the panoramic image, the correspondence between the detail images of each region on the panoramic image and the correspondence between the detail images of each region can be determined, thereby achieving the stitching of the detail images of each region.
[0024] In this embodiment, a single array camera can monitor terminal screens within a defined monitoring area. This area can contain one or more terminals. A single terminal screen can display the monitoring video corresponding to one monitored image, or it can display the monitoring videos corresponding to multiple monitored images. The array camera includes one wide-angle lens and n detail lenses, resulting in an image set of a rows and b columns. n is an integer not less than 1, and the user can set n according to the monitoring level of the abnormal event. The number of n is proportional to the monitoring level; that is, the higher the monitoring level, the more detail lenses are used.
[0025] The wide-angle lens captures a panoramic image of the surveillance video displayed on the terminal screen, while the detail lens captures detailed images of various areas within the surveillance video displayed on the terminal screen.
[0026] Step S20: Based on the target matching point set corresponding to the panoramic image and each regional detail image, calculate the transformation matrix corresponding to each regional detail image on the panoramic image; In this embodiment, moving objects in the surveillance video are tracked, and the corresponding points of the same moving object at different times in the overlapping parts of various regional detail images are used as matching points. A homography matrix is then calculated based on these matching points. This avoids the problem of insufficient and inaccurate matching feature points often found in static images, which can lead to errors in perspective transformation matrix calculation, or the problem of feature matching deviations caused by too many repetitive textures in the image, thereby reducing stitching errors.
[0027] Specifically, the target matching point set based on the panoramic image and each regional detail image includes: Upon receiving a detection instruction for a target anomalous event, a feature detector is determined in either an accelerated robust feature converter or a scale-invariant feature converter based on the event type of the target anomalous event, so as to generate the target matching point set based on the feature detector; Upon receiving a switching instruction for a detected event, the feature detector is switched between an accelerated robust feature converter and a scale-invariant feature converter based on the type of the event after the switch, so as to generate the target matching point set based on the switched feature detector.
[0028] In this embodiment, users can perform targeted detection of specific abnormal events in the surveillance video according to actual application needs. For example, they can perform targeted detection of high-precision abnormal event types such as abnormal gathering or fire in the surveillance video, and can also perform targeted detection of abnormal event types with high real-time requirements such as bird detection during flight or foreign object detection during autopilot.
[0029] Furthermore, when the target abnormal event monitoring period ends or the type of the detected target abnormal event changes, the current accelerated robust feature converter or scale-invariant feature converter can be switched in real time according to the changed event type.
[0030] Specifically, replaceable feature detectors can be created for image feature matching. Different types of feature detectors can be used depending on the required feature point accuracy and matching speed. For example, for monitoring scenarios with high accuracy requirements (such as object abandonment and target personnel search), SIFT (Scale-invariant feature transform) can be used as the feature detector; for monitoring scenarios with high real-time requirements (such as bird detection during aircraft flight and fire detection), SURF (Speeded UpRobust Features) can be used. Thus, feature detectors can be switched according to different monitoring scenarios. Furthermore, the feature detector can be switched accordingly when the target abnormal event concludes or the type of target abnormal event changes, improving the matching degree between the feature detector and the monitoring scenario and thus meeting the monitoring needs of different scenarios.
[0031] In a specific embodiment, a feature detector can also be used to perform feature detection on the panoramic image and the regional detail image respectively, and calculate the feature points corresponding to the panoramic image and the regional detail image. For example, the k-nearest neighbor matching (knnmatch) algorithm can be used to perform feature matching on the feature points of the panoramic image and the regional detail image to obtain the feature points after feature matching. Then, the feature points after feature matching are filtered to ensure that the distance between the feature points is within a preset range, thereby removing mismatched points and obtaining the target matching point set. After obtaining the target matching point set, the homography matrix (Homography matrix) is calculated using a robust method based on RANSAC to obtain the transformation matrix T (i.e., the scale transformation matrix).
[0032] Step S30: Based on each transformation matrix, stitch together the detailed images of each region to obtain a stitched image, which serves as the detailed magnified image corresponding to the panoramic image. Then, based on the detailed magnified image, detect whether there are any abnormal events in the monitoring video.
[0033] In this embodiment, a perspective transformation is performed on each of the aforementioned regional detail images according to the scale transformation matrix T corresponding to each image, resulting in a perspective-transformed image. The surrounding pixels of each transformed regional detail image are cropped (i.e., the black borders caused by the transformation are removed), thus obtaining n magnified regional detail images. These transformed regional detail images are then sequentially placed on their corresponding regions on the panoramic image, and stitched together to obtain the magnified detail image corresponding to the panoramic image. Therefore, based on the magnified detail image corresponding to the panoramic image, accurate identification of whether an abnormal event has occurred in the surveillance video can be performed.
[0034] This embodiment provides a surveillance video detection method based on an array camera. The method uses an array camera to acquire surveillance video from at least one terminal screen, obtaining a panoramic image and at least one regional detail image of the surveillance video. Based on the target matching point set corresponding to the panoramic image and each regional detail image, a transformation matrix corresponding to each regional detail image on the panoramic image is calculated. Based on each transformation matrix, the regional detail images are stitched together to obtain a stitched image, which serves as a magnified detail image corresponding to the panoramic image. The presence of abnormal events in the surveillance video is then detected based on the magnified detail image. Through this method, this application monitors surveillance video using an external camera outside the terminal screen, capturing the surveillance video displayed on the terminal screen using an array camera. By magnifying the details of each frame of the surveillance video using the panoramic image and regional detail images, the image resolution of the surveillance video is improved. By identifying abnormal events in the high-resolution image, the accuracy of abnormal event detection in the surveillance video is improved. Therefore, abnormal detection of surveillance video on any terminal can be achieved without installing any software on the terminal, improving the convenience of surveillance video monitoring.
[0035] Furthermore, to adapt to different monitoring scenarios (such as high monitoring accuracy or fast response time), a replaceable feature detector is created, providing two feature detectors, SIFT and SURF, to provide image feature matching under different feature matching requirements.
[0036] Specifically, preprocessing operations such as noise reduction and brightness adjustment are performed on the surveillance video. After preprocessing, panoramic images and detailed images of each area are acquired from the surveillance video.
[0037] Based on the static feature matching strategy, feature point detection and matching algorithms, such as SIFT, SURF, or RANSAC (Random Sample Consensus), are used for each pair of adjacent frames in the panoramic image and each regional detail image to extract the initial set of matching points, which is then used as the static matching point set.
[0038] To avoid the problem of insufficient and inaccurate matching feature points in static images, which often leads to errors in the calculation of the perspective transformation matrix, we further use moving object detection and tracking to find more accurate matching points in the panoramic image and the detailed images of each region to calculate the Homography matrix.
[0039] Specifically, motion detection methods, such as frame difference, optical flow, or object detection, such as YOLO, are first used to detect moving objects in each frame. Then, tracking algorithms, such as optical flow tracing or SORT (Simple Online and Realtime Tracking, a multi-object tracking algorithm), are used to obtain the trajectory of the moving object in the frame sequence. Feature points in the overlapping regions are then extracted based on the trajectory of the moving object to construct a dynamic matching point set.
[0040] In a specific embodiment, erroneous matching points in the target matching point set, composed of static and dynamic matching point sets, can be removed based on RANSAC. Alternatively, erroneous matching points can be removed based on the distance between each matching point in the target matching point set.
[0041] By combining static and dynamic matching point sets, the Homography matrices corresponding to the panoramic image and each regional detail image are calculated. These calculated Homography matrices transform each regional detail image to a unified coordinate system similar to the panoramic image, facilitating the stitching of these images. The stitched boundaries are then optimized (e.g., seamless blending or gradual transition). The video sequence is processed frame-by-frame to generate a complete stitched video or image. Finally, further enhancements (e.g., color balancing, ghosting removal) can be applied to improve the overall image quality.
[0042] Furthermore, the step of stitching together the detailed images of each region based on the various transformation matrices to obtain a stitched image includes: Based on the transformation matrices corresponding to each of the aforementioned regional detail images, the rectangular region of each regional detail image on the panoramic image is calculated. In each rectangular region of the panoramic image, the largest bounding rectangle region is determined, and the panoramic image is cropped based on the largest bounding rectangle region. The cropped panoramic image is enlarged, and feature point matching is performed on each of the region detail images and the cropped and enlarged panoramic image to obtain the transformation matrix of each region detail image on the cropped and enlarged panoramic image. Based on the transformation matrix of each of the aforementioned regional detail images on the cropped and enlarged panoramic image, the aforementioned regional detail images are stitched together to obtain a stitched image.
[0043] Specifically, the step of stitching together the detailed images of each region based on the various transformation matrices to obtain the stitched image further includes: Based on the transformation matrices corresponding to each of the aforementioned regional detail images, perspective transformation is performed on four points of each of the aforementioned regional detail images to obtain four points of each regional detail image on the panoramic image. For each regional detail image, a rectangle is formed by connecting four points on the panoramic image to obtain the bounding box of each regional detail image on the panoramic image; Within the mapping frame of each regional detail image on the panoramic image, the minimum width and height and the maximum width and height are determined, and the scaling factor corresponding to each regional detail image is calculated based on the target width and height, the minimum width and height and the maximum width and height. Based on the scaling factor corresponding to each regional detail image, the size of each regional detail image is adjusted, and the adjusted regional detail images are stitched together to form the stitched image.
[0044] In this embodiment, after obtaining the transformation matrix T corresponding to the panoramic image and each regional detail image, perspective transformation is performed on the four points of each regional detail image according to the transformation matrix T to obtain the four points of each regional detail image on the panoramic image, and the bounding rectangle of these four points is calculated in the panoramic image to obtain the mapping box of each regional detail image on the panoramic image.
[0045] Specifically, the bounding matrix of each region detail image in the panoramic image is calculated to obtain the boundRect of all region detail images in the panoramic image. In the panoramic image, the regions corresponding to the boundRect of all region sub-images in the panoramic image are cropped and enlarged. The enlargement formula is (3840*a+200, 2160*b+200); the value of +200 can be adjusted appropriately according to the size of the black border after the actual detail image transformation, thus reserving space for subsequent black border cropping, thereby achieving a resolution of (3840*a, 2160*b) in the stitched result image. After enlarging the panoramic image, feature detection and matching are performed again with each detail image to calculate the corresponding transformation matrix. The resulting transformation matrix is the enlarged version, and the transformed detail image also has the enlarged resolution. That is: First, a feature matching process is performed between the panoramic image and the detailed images of each region: Feature point matching is performed on each of the detailed regional images (i.e., n sub-images) and the panoramic image to obtain the panoramic / sub-image transformation matrix T corresponding to each sub-image on the panoramic image. Then, the rectangular region of each sub-image on the panoramic image is calculated using the transformation matrix T. Based on the rectangular regions of the n sub-images, the largest enclosing rectangular region is found, where the largest enclosing rectangular region includes all the rectangular regions of the n sub-images. The panoramic / sub-image transformation matrix T, the rectangular region rect, and the largest enclosing rectangular region are recorded for each sub-image.
[0046] Then, the rectangular areas containing detailed images of each region in the panoramic image are magnified: The panoramic image is cropped based on the largest enclosing rectangular area to obtain the cropped panoramic image. The cropped image is then enlarged to (e.g., 1920*b+200, 1080*a+200). By enlarging the cropped panoramic image by (+200), the stitched image avoids exceeding the edges when matching the panoramic image with the detailed images of each region in the subsequent stitching process. This reserve for the subsequent stitching results' cropping edges improves the resolution of the final stitched image (e.g., 1920*b, 1080*a).
[0047] Then, a second feature matching process is performed between the cropped panoramic image and the detailed images of each region: Feature point matching is performed on each of the n sub-images and the cropped and enlarged panoramic image. The scale transformation matrix T_scale of each sub-image on the panoramic image is calculated. The scale rectangular region rect_scale of each sub-image on the panoramic image is then calculated using the scale transformation matrix T_scale. The scale transformation matrix T_scale of each sub-image and the panoramic image, as well as the scale rectangular region rect_scale, are recorded.
[0048] In further embodiments, to avoid enlarging the panoramic image and recalculating feature detection, matching, and transformation matrices, especially since feature detection, matching, and transformation matrix calculations are very time-consuming at high resolutions, a scaling factor is calculated based on the region of each detail image within the panoramic image to further reduce the computation time of the transformation matrix. The scaling factor is calculated as follows: Find the minimum width and height within all region boxes (ensuring the minimum can be scaled to the specified resolution (3840, 2160), while larger ones can be cropped). The scaling factor is derived by calculating the scale ratio based on the minimum width and height and (target width and height + edge increment). The edge increment is used for subsequent black border cropping after transformation, ensuring that each detail image reaches the target resolution (3840, 2160) after cropping. The formula for calculating the scaling factor is: Sfactor=( min(rect.witdh) / (Wtarget+△W), min(rect.height) / (Htarget+△H) ).
[0049] After calculating the scale factor, the scaling factor parameter in the transformation matrix is replaced with the scale factor. The detail image is then transformed using the transformation matrix to obtain the magnified transformation result.
[0050] The above method involves multiple steps, including processing image feature points, calculating the homography matrix, matching bounding boxes, and image scaling, to complete the preparatory operations before image stitching. The adjusted stitching parameters generated during this process are then saved to a calibration file. Specifically: Feature point extraction and matching: For the baseline image (i.e., the panoramic image) and each frame of detail images (i.e., each region's detail images), feature points and descriptors are extracted and then matched. Successfully matched feature points are used to calculate the homography matrix.
[0051] Calculate the homography matrix: Use feature matching to calculate the homography matrix between the baseline image and the detail image.
[0052] Bounding box processing: Calculate matching bounding boxes for each image. Calculate the enclosing rectangle of all bounding boxes as the overall layout for stitching the images.
[0053] Image scaling: Images are adjusted by calculating a scaling factor so that they fit into the stitched panorama. The scaling factor is calculated based on the minimum size of the matching rectangle.
[0054] Display Rectangle: Calculate the final display position based on the scaled rectangle and calculate a suitable display area for the stitched image.
[0055] Saving stitching parameters: The final generated stitching parameters (including homography matrix, bounding box, scaling factor, etc.) are saved to a calibration file for subsequent image stitching. For example, the panorama / sub-image transformation matrix T, the bounding box rect, and the largest enclosing bounding box recorded from the first feature matching, and the scale transformation matrix T_scale and the scale bounding box rect_scale recorded from the second feature matching are recorded as calibration parameters in the calibration file. Subsequently, the calibration parameters can be applied to obtain the transformed sub-images. That is, based on the scale transformation matrix T of each sub-image, each sub-image undergoes perspective transformation to obtain the perspective-transformed image. Each transformed sub-image is then cropped by removing surrounding pixels (removing the black borders caused by the transformation), resulting in the final n transformed sub-images.
[0056] Furthermore, the step of detecting whether there are abnormal events in the surveillance video based on the magnified detail image specifically includes: Obtain the abnormal event features corresponding to each abnormal event in the abnormal event list, and perform abnormal event detection on the magnified detail image based on each abnormal event feature; If an intrusion event is detected in the magnified detailed image, the intrusion target is marked in the monitoring video of all terminal screens; The array camera records the movement trajectory of the intrusion target in the surveillance video, and based on the intrusion target and its corresponding movement trajectory, generates and displays abnormal early warning information of the intrusion event.
[0057] In this embodiment, sub-images captured by an array camera are matched against a panoramic image to determine the position of each sub-image. The sub-images are then stitched together according to their positions, and image frames for each sub-image are displayed on the stitched image, with colors assigned to the frames. For example, a green frame indicates no anomaly, a red frame indicates an anomaly, and a yellow frame indicates a warning, providing visual feedback and improving user experience. A user-preset list of abnormal events or warning events is obtained. Based on this list, the stitched image is identified, such as abnormal behavior (e.g., running, gathering), intrusion, fire, equipment damage, or items left behind. When an intrusion is detected, the same intruder is associated across multiple sub-images, monitoring the intruder's position. As the intruder moves from one array camera view to another, the color of the image frame changes to indicate the intruder's current location, continuing until the intruder leaves all viewpoints, thus obtaining surveillance video of the intruder. The obtained surveillance video is analyzed to determine if the intruder exhibits any abnormal or warning-required behavior.
[0058] Specifically, user-preset event priorities can be obtained. When an abnormal event or warning event is detected, an abnormal warning scheme (such as a red and flashing image box, alarm sound, alarm duration, etc.) is determined according to the priority of the abnormal behavior, and the abnormal handling personnel are notified and the emergency plan is activated.
[0059] To ensure the security of surveillance video, existing anomaly detection solutions only allow manual monitoring. Therefore, are there corresponding security measures for monitoring computer screens captured by array cameras to prevent the leakage of the captured content? It is possible to encrypt and save the content captured by the array cameras for later review.
[0060] For indoor monitoring scenarios, a variable focus and variable spacing design is adopted to adapt to different indoor scenes. A high refresh rate lens is used to avoid ripple imaging when sampling the screen. As a result, the array camera in this application can adapt to the deployment of various indoor monitoring scenarios, making it easy to see all the monitoring screens clearly and comprehensively, and facilitating subsequent automatic identification.
[0061] Specifically, this embodiment further provides a multi-screen monitoring system based on a zoom array camera, the system comprising: The variable-focus array camera module consists of multiple camera units, supporting dynamic adjustment of focal length (5-50mm), field of view (e.g., 60°-120°), and camera spacing (10-50cm), and is equipped with a high refresh rate lens (≥120Hz) to eliminate screen ripples; The intelligent detection module includes a color block recognition unit, a text status recognition unit, and a dynamic interference filtering unit, which are used to analyze the sensor status information on the monitoring screen and eliminate environmental interference. The user configuration module provides a graphical interface for users to annotate screen monitoring areas and define alert rules; To address the issue of different monitoring components using different colors on the monitoring interface, a dedicated universal recognition module, such as color block recognition, is built. To facilitate the recognition of status text on the screen, a visual recognition model, a status box recognition module, and a general object recognition module can be specifically trained. Furthermore, a user-defined interface for the desired recognition is provided, allowing users to flexibly customize it according to their actual needs and define monitoring trigger states themselves without the need for developers.
[0062] The early warning output module triggers audible and visual alarms or pop-up notifications based on the detection results.
[0063] The variable focal length array camera module includes: The low-light enhancement unit adapts to complex lighting environments through multi-frame noise reduction and dynamic exposure compensation; The wide-angle distortion correction unit uses a fisheye lens combined with a perspective transformation algorithm to ensure complete imaging at the screen edges.
[0064] The color block recognition unit: Based on HSV color space threshold segmentation, preset color regions (such as red warning, green normal) are extracted from the screen. By combining the coordinates of the fixed area calibrated during installation, interference from similar colors in the environment is filtered out.
[0065] The text state recognition unit includes: A deep learning-based OCR model (CRNN+Attention) can be used to recognize status text such as "fault" and "normal" on the screen. The semantic analysis module associates text with sensor values (such as temperature and pressure) to identify anomalies.
[0066] The dynamic interference filtering unit includes: During the installation phase, the positions of screen controls are marked, and a static background template is generated. Background difference method is used to detect dynamic interference (such as people walking), and time continuity analysis is combined to eliminate instantaneous interference.
[0067] To prevent the color block recognition module from being easily interfered with by color blocks in the environment, or by objects of similar color or workers appearing in the frame, a restricted area definition is added to the algorithm to reduce false detections. Since the array camera is fixed after installation, the monitored screen area and the positions of its controls are known and unchanging. The location area information for each monitoring state is calibrated during the configuration process. The algorithm to reduce false detections uses sufficient information to determine whether temporarily appearing interfering color blocks in the frame belong to the category requiring an alert. This same false detection prevention mechanism can also be used for the misidentification of other elements, such as text status recognition, or for monitoring screens that do not have a built-in AI recognition module and require the general object recognition module in this monitoring system to identify intruders, vehicles, etc.
[0068] The user configuration module supports: Drag-and-drop annotation tool to define monitoring areas (such as temperature display box, carbon monoxide alarm light); Configure logical rules and set composite warning conditions (such as a red block lasting for 5 seconds and a pressure value > 100 kPa).
[0069] Based on the aforementioned system, this application enables monitoring of multiple screens indoors, specifically: Deploy variable-focus array cameras, adjusting the focal length and spacing to cover all monitoring screens; The screen control area and warning rules are defined through the user configuration module; The system captures screen images in real time, and outputs anomaly warnings after color block recognition, text detection, and interference filtering.
[0070] The interference filtering step is as follows: The detection range is limited by the calibration area, and only the position of screen controls is analyzed. Confidence weights are applied to color blocks or text in non-fixed areas (such as environmental backgrounds), and those below a threshold are considered interference.
[0071] Addressing the issue of low efficiency in current manual monitoring, this application employs a variable-focus, variable-pitch array camera. Through adaptive focal length (5-50mm), field of view (60°-120°), and a high refresh rate lens (≥120Hz) to resist screen ripples, it monitors the status of various sensors on the monitoring screen, adapting to indoor multi-screen monitoring scenarios. Combined with color block recognition, text status detection, and dynamic interference filtering algorithms, it achieves automated anomaly detection of sensor status (such as temperature, pressure, carbon monoxide, etc.) on the screen. A no-code configuration interface is provided, supporting user-defined monitoring rules. This application does not require integration with existing monitoring systems, reduces security risks through non-intrusive deployment, and significantly improves monitoring efficiency (reducing the missed detection rate by over 90%), making it suitable for indoor multi-screen monitoring scenarios such as airports and power plants.
[0072] Please see Figure 2 , Figure 2 This application provides a schematic diagram of the functional modules of a surveillance video detection device based on an array camera.
[0073] like Figure 2 As shown, the surveillance video detection device 200 based on an array camera includes: The terminal screen monitoring module 210 is used to acquire monitoring video from at least one terminal screen through an array camera, and to acquire a panoramic image of the monitoring video and at least one regional detail image. The transformation matrix generation module 220 is used to calculate the transformation matrix corresponding to each regional detail image on the panoramic image based on the target matching point set corresponding to the panoramic image and each regional detail image. The abnormal event detection module 230 is used to stitch together the detailed images of each region based on each transformation matrix to obtain a stitched image, which serves as a magnified detail image corresponding to the panoramic image, and to detect whether there are abnormal events in the monitoring video based on the magnified detail image.
[0074] Furthermore, the target matching point set includes a static matching point set and a dynamic matching point set, and the surveillance video detection device 200 based on the array camera further includes: The static matching point acquisition module is used to perform feature point detection and feature point matching on each pair of adjacent frames in the monitoring video based on a replaceable feature detector to obtain the static matching point set. A moving object detection module is used to detect moving objects in each frame of the surveillance video based on a preset moving object detection module. The moving object detection module includes a frame difference detection module, an optical flow detection module, and a YOLO detection module. A motion trajectory tracking module is used to acquire the motion trajectory of the moving object in a frame image sequence based on a preset tracking module, wherein the tracking module includes an optical flow tracking module and a multi-target tracking module; The dynamic matching point acquisition module is used to extract feature points in the overlapping areas of the panoramic image and each regional detail image based on the motion trajectory of the moving object, and use them as the dynamic matching point set.
[0075] Furthermore, the surveillance video detection device 200 based on an array camera also includes a feature detector switching module, used for: When the feature detection accuracy of the feature detector is greater than the preset accuracy threshold, an accelerated and robust feature converter is used as the alternative feature detector to accurately detect matching points in monitoring images with high feature accuracy requirements. When the feature detection speed of the feature detector exceeds a preset speed threshold, a scale-invariant feature converter is used as the alternative feature detector to quickly detect matching points in monitoring images with high response time requirements.
[0076] Furthermore, the abnormal event detection module specifically includes: The vertex detection unit is used to perform perspective transformation on four points of each of the aforementioned regional detail images based on the respective transformation matrices corresponding to each of the regional detail images, so as to obtain four points of each regional detail image on the panoramic image. The mapping frame determination unit is used to calculate the rectangle connecting four points of each regional detail image on the panoramic image to obtain the mapping frame of each regional detail image on the panoramic image. The scaling factor determination unit is used to determine the minimum width and height and the maximum width and height within the mapping frame of each regional detail image on the panoramic image, and to calculate the scaling factor corresponding to each regional detail image based on the target width and height, the minimum width and height and the maximum width and height. The regional image stitching unit is used to adjust the size of each regional detail image based on the scaling factor corresponding to each regional detail image, and stitch the adjusted regional detail images into the stitched image.
[0077] Furthermore, the abnormal event detection module specifically includes: The rectangular region determination unit is used to calculate the rectangular region of each regional detail image on the panoramic image based on each transformation matrix corresponding to each of the regional detail images; A panoramic image cropping unit is used to determine the largest enclosing rectangular region among various rectangular regions on the panoramic image, and to crop the panoramic image based on the largest enclosing rectangular region. The transformation matrix calculation unit is used to enlarge the cropped panoramic image and perform feature point matching on each of the regional detail images and the cropped and enlarged panoramic image to obtain the transformation matrix of each of the regional detail images on the cropped and enlarged panoramic image. The image stitching generation unit is used to stitch together the various regional detail images based on the transformation matrix of each regional detail image on the cropped and enlarged panoramic image to obtain a stitched image.
[0078] Furthermore, the transformation matrix generation module specifically includes: The feature detector determination unit is used to determine a feature detector in an accelerated robust feature converter or a scale-invariant feature converter based on the event type of the target abnormal event when a detection instruction for a target abnormal event is received, so as to generate the target matching point set based on the feature detector; The feature detector switching unit is used to switch the feature detector between an accelerated robust feature converter or a scale-invariant feature converter based on the event type after receiving a switching instruction for a detection event, so as to generate the target matching point set based on the switched feature detector.
[0079] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and modules can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0080] The aforementioned device can be implemented as a computer program, which can be used in, for example... Figure 3 It runs on the computer device shown.
[0081] Please see Figure 3 , Figure 3 This is a schematic block diagram illustrating the structure of a computer device according to an embodiment of this application. The computer device may be a server.
[0082] See Figure 3 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0083] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any array-based camera-based surveillance video detection method.
[0084] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0085] Internal memory provides an environment for the execution of computer programs stored in non-volatile storage media. When executed by a processor, the computer program enables the processor to perform any surveillance video detection method based on an array camera.
[0086] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0087] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0088] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: In one embodiment, the processor is configured to run a computer program stored in memory and also to implement: Using an array camera, acquire surveillance video from at least one terminal screen, and acquire a panoramic image and at least one area detail image of the surveillance video; Based on the target matching point set corresponding to the panoramic image and each regional detail image, calculate the transformation matrix corresponding to each regional detail image on the panoramic image; Based on each transformation matrix, the detailed images of each region are stitched together to obtain a stitched image, which serves as the detailed magnified image corresponding to the panoramic image. The presence of abnormal events in the surveillance video is then detected based on the detailed magnified image.
[0089] In one embodiment, the processor is configured to run a computer program stored in memory and also to implement: Based on a replaceable feature detector, feature point detection and feature point matching are performed on each pair of adjacent frames in the surveillance video to obtain the static matching point set. Based on a preset moving object detection module, moving objects in each frame of the surveillance video are detected. The moving object detection module includes a frame difference detection module, an optical flow detection module, and a YOLO detection module. Based on a preset tracking module, the motion trajectory of the moving object in the frame image sequence is obtained, wherein the tracking module includes an optical flow tracking module and a multi-target tracking module; Based on the motion trajectory of the moving object, feature points in the overlapping areas of the panoramic image and each regional detail image are extracted as the dynamic matching point set.
[0090] In one embodiment, the processor is configured to run a computer program stored in memory and also to implement: When the feature detection accuracy of the feature detector is greater than the preset accuracy threshold, an accelerated and robust feature converter is used as the alternative feature detector to accurately detect matching points in monitoring images with high feature accuracy requirements. When the feature detection speed of the feature detector exceeds a preset speed threshold, a scale-invariant feature converter is used as the alternative feature detector to quickly detect matching points in monitoring images with high response time requirements.
[0091] In one embodiment, the processor is configured to run a computer program stored in memory and also to implement: Based on the transformation matrices corresponding to each of the aforementioned regional detail images, perspective transformation is performed on four points of each of the aforementioned regional detail images to obtain four points of each regional detail image on the panoramic image. For each regional detail image, a rectangle is formed by connecting four points on the panoramic image to obtain the bounding box of each regional detail image on the panoramic image; Within the mapping frame of each regional detail image on the panoramic image, the minimum width and height and the maximum width and height are determined, and the scaling factor corresponding to each regional detail image is calculated based on the target width and height, the minimum width and height and the maximum width and height. Based on the scaling factor corresponding to each regional detail image, the size of each regional detail image is adjusted, and the adjusted regional detail images are stitched together to form the stitched image.
[0092] In one embodiment, the processor is configured to run a computer program stored in memory and also to implement: Based on the transformation matrices corresponding to each of the aforementioned regional detail images, the rectangular region of each regional detail image on the panoramic image is calculated. In each rectangular region of the panoramic image, the largest bounding rectangle region is determined, and the panoramic image is cropped based on the largest bounding rectangle region. The cropped panoramic image is enlarged, and feature point matching is performed on each of the region detail images and the cropped and enlarged panoramic image to obtain the transformation matrix of each region detail image on the cropped and enlarged panoramic image. Based on the transformation matrix of each of the aforementioned regional detail images on the cropped and enlarged panoramic image, the aforementioned regional detail images are stitched together to obtain a stitched image.
[0093] In one embodiment, the processor is configured to run a computer program stored in memory and also to implement: Upon receiving a detection instruction for a target anomalous event, a feature detector is determined in either an accelerated robust feature converter or a scale-invariant feature converter based on the event type of the target anomalous event, so as to generate the target matching point set based on the feature detector; Upon receiving a switching instruction for a detected event, the feature detector is switched between an accelerated robust feature converter and a scale-invariant feature converter based on the type of the event after the switch, so as to generate the target matching point set based on the switched feature detector.
[0094] In one embodiment, the processor is configured to run a computer program stored in memory and also to implement: Obtain the abnormal event features corresponding to each abnormal event in the abnormal event list, and perform abnormal event detection on the magnified detail image based on each abnormal event feature; If an intrusion event is detected in the magnified detailed image, the intrusion target is marked in the monitoring video of all terminal screens; The array camera records the movement trajectory of the intrusion target in the surveillance video, and based on the intrusion target and its corresponding movement trajectory, generates and displays abnormal early warning information of the intrusion event.
[0095] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the surveillance video detection methods based on array cameras provided in the embodiments of this application.
[0096] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.
[0097] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for detecting surveillance video based on an array camera, characterized in that, The array camera includes one wide-angle lens and n detail lenses, and the surveillance video detection method includes the following steps: The number of events n is determined based on the monitoring level of the abnormal events to be detected, and the number of events n is proportional to the monitoring level. The panoramic image is captured through the wide-angle lens, and the detailed image of the region is captured through the n detail lenses, where n is an integer not less than 1; Feature detection is performed on the panoramic image and each regional detail image respectively, and the feature points corresponding to the panoramic image and each regional detail image are calculated as the target matching point set corresponding to the panoramic image and each regional detail image. The homography matrix corresponding to the target matching point set is calculated as the transformation matrix corresponding to each regional detail image on the panoramic image. Based on each transformation matrix, the detailed images of each region are stitched together to obtain a stitched image, which serves as the detailed magnified image corresponding to the panoramic image. The presence of abnormal events in the surveillance video is then detected based on the detailed magnified image.
2. The surveillance video detection method based on an array camera as described in claim 1, characterized in that, The step of performing feature detection on the panoramic image and each regional detail image, calculating the feature points corresponding to the panoramic image and each regional detail image as the target matching point set corresponding to the panoramic image and each regional detail image, and calculating the homography matrix corresponding to the target matching point set as the transformation matrix corresponding to each regional detail image on the panoramic image, includes: The feature points of the panoramic image and the regional detail image are matched using the k-nearest neighbor matching algorithm to obtain the feature points after feature matching. Based on the distance between feature points, the feature points after feature matching are filtered to filter out mismatched points whose distance between feature points is outside a preset range, thereby obtaining the target matching point set; A robust RANSAC-based method is used to calculate the homography matrix corresponding to the target matching point set, which is then used as the transformation matrix.
3. The surveillance video detection method based on an array camera as described in claim 1, characterized in that, The abnormal events include a first event type requiring high matching accuracy and a second event type requiring high matching speed. Before tracking moving objects in the surveillance video and determining the target matching point set based on the corresponding points of the same moving object at different times in the overlapping parts of each regional detail image, and calculating the homography matrix corresponding to the target matching point set as the transformation matrix corresponding to each regional detail image on the panoramic image, the process further includes: When the type of the abnormal event is the first event type, the scale-invariant feature transformer SIFT is used as a replacement feature detector; When the type of the abnormal event is the second time type, the accelerated robust feature converter SURF is used as an alternative feature detector. Based on a replaceable feature detector, feature detection is performed on the panoramic image and each regional detail image respectively, and the feature points corresponding to the panoramic image and each regional detail image are calculated as the target matching point set corresponding to the panoramic image and each regional detail image. The homography matrix corresponding to the target matching point set is calculated as the transformation matrix corresponding to each regional detail image on the panoramic image.
4. The surveillance video detection method based on an array camera as described in claim 1, characterized in that, The step of stitching together the detailed images of each region based on various transformation matrices to obtain a stitched image, which serves as the detailed magnified image corresponding to the panoramic image, includes: Based on the transformation matrix corresponding to each of the aforementioned regional detail images, a perspective transformation is performed on each of the aforementioned regional detail images to obtain a perspective-transformed image; By cropping the surrounding pixels from each transformed area detail image, a magnified area detail image is obtained. The magnified detail images of the regions are placed sequentially on the corresponding regions of the panoramic image, and then stitched together to obtain the magnified detail images corresponding to the panoramic image.
5. The surveillance video detection method based on an array camera as described in claim 1, characterized in that, The method further includes: Obtain the stitching parameters and save them to a calibration file to perform detailed image stitching and / or panoramic image detail magnification based on the stitching parameters in the calibration file; The stitching parameters include at least one of the following: feature points corresponding to the panoramic image and each regional detail image, homography matrix, rectangular bounding boxes of each regional detail image on the panoramic image, or scaling factors.
6. The surveillance video detection method based on an array camera as described in claim 1, characterized in that, The detection of abnormal events in the surveillance video based on the magnified detail image specifically includes: Obtain the abnormal event features corresponding to each abnormal event in the abnormal event list. Based on each abnormal event feature, perform abnormal event detection on the magnified detailed image and display the abnormal event detection results of each priority through image frames of different colors. The color block recognition unit monitors each color region in the magnified detail image based on HSV color space threshold segmentation.
7. The surveillance video detection method based on an array camera as described in any one of claims 1-6, characterized in that, The target matching point set includes a static matching point set and a dynamic matching point set. Before performing feature detection on the panoramic image and each regional detail image to calculate the feature points corresponding to the panoramic image and each regional detail image, and using these as the target matching point set corresponding to the panoramic image and each regional detail image, the method further includes: Based on a replaceable feature detector, feature point detection and feature point matching are performed on each pair of adjacent frames in the surveillance video to obtain the static matching point set. Based on a preset moving object detection module, moving objects in each frame of the surveillance video are detected. The moving object detection module includes a frame difference detection module, an optical flow detection module, and a YOLO detection module. Based on a preset tracking module, the motion trajectory of the moving object in the frame image sequence is obtained, wherein the tracking module includes an optical flow tracking module and a multi-target tracking module; Based on the motion trajectory of the moving object, feature points in the overlapping areas of the panoramic image and each regional detail image are extracted as the dynamic matching point set.
8. A surveillance video detection device based on an array camera, characterized in that, The array camera includes one wide-angle lens and n detail lenses, and the surveillance video detection device based on the array camera includes: The detail shot determination module is used to determine n based on the monitoring level of the abnormal event to be detected, and the number of n is proportional to the monitoring level. The terminal screen monitoring module is used to acquire the panoramic image through the wide-angle lens and to acquire the regional detail image through the n detail lenses, where n is an integer not less than 1; The transformation matrix calculation module is used to perform feature detection on the panoramic image and each regional detail image respectively, calculate the feature points corresponding to the panoramic image and each regional detail image as the target matching point set corresponding to the panoramic image and each regional detail image, and calculate the homography matrix corresponding to the target matching point set as the transformation matrix corresponding to each regional detail image on the panoramic image. The abnormal event monitoring module is used to stitch together the detailed images of each region based on various transformation matrices to obtain a stitched image, which serves as a magnified detail image corresponding to the panoramic image, and to detect whether there are abnormal events in the monitoring video based on the magnified detail image.
9. A computer device, characterized in that, The computer device includes a processor, a memory, and a camera-based surveillance video detection program stored in the memory and executable by the processor, wherein when the camera-based surveillance video detection program is executed by the processor, it implements the steps of the camera-based surveillance video detection method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a surveillance video detection program based on an array camera, wherein when the surveillance video detection program based on an array camera is executed by a processor, it implements the steps of the surveillance video detection method based on an array camera as described in any one of claims 1 to 7.