System and method for load balancing in video surveillance system
By implementing load balancing among cameras in a video surveillance system, and by identifying areas of interest and sending them to other cameras for video analysis, the problem of bandwidth limitations is solved, thereby improving the processing power and efficiency of video analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2026-03-27
AI Technical Summary
In video surveillance systems, the limited processing bandwidth of cameras leads to unbalanced loads, affecting the simultaneous operation capability of video analysis algorithms.
By load balancing among multiple networked cameras, the first camera identifies the region of interest and sends it to the second camera for video analysis. The second camera performs the video analysis and returns the results.
It achieves load balancing among cameras, improves the processing power and efficiency of video analysis, and reduces processing latency.
Smart Images

Figure CN121750818A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates generally to video surveillance systems, and more specifically to balanced workloads within video surveillance systems. Background Technology
[0002] A video surveillance system may include a large number of cameras, each of which generates a video stream. The cameras may include various video analysis algorithms that the cameras can use when analyzing the video streams captured by those cameras. In some cases, it may be desirable to execute multiple video analysis algorithms on a particular video stream or even on a specific video frame within a particular video stream. Cameras may have processing bandwidth limitations, which may affect how many video analysis algorithms a camera can run simultaneously. Meanwhile, other cameras within the video surveillance system may have available processing bandwidth due to, for example, a lack of activity detected in the video streams of other cameras. Systems and methods for load balancing among cameras within a video surveillance system are desired. Systems and methods that allow a first camera to request assistance from a second camera to run one or more specified video analysis algorithms that the first camera currently lacks the available processing bandwidth for. Summary of the Invention
[0003] This disclosure relates generally to video surveillance systems, and more specifically to load balancing within video surveillance systems. An example may exist in a method for load balancing video analytics processing among two or more networked cameras in a plurality of networked cameras, wherein each networked camera includes a camera for capturing a corresponding video stream. This exemplary method includes a first camera in the plurality of networked cameras identifying an object of interest in a video frame of a video stream captured by the first camera. A region of interest (ROI) corresponding to the object of interest is cropped from the video frame of the video stream, wherein the cropped region includes less than the entire video frame of the video stream. The cropped ROI of the video frame of the video stream is then sent to a second camera in the plurality of networked cameras, and the second camera in the plurality of networked cameras performs a video analytics algorithm on the cropped ROI of the video frame of the video stream to produce a video analytics result. The second camera in the plurality of networked cameras sends the video analytics result to the first camera in the plurality of networked cameras, a network video recorder, and / or another device.
[0004] Another example could exist in a surveillance system. This system includes a first camera, a network, and a second camera, the second camera being operatively coupled to the first camera via the network. The first camera is configured to capture and process a video stream to identify regions of interest (ROIs) in video frames corresponding to motion within those frames, identify objects of interest (ROIs) corresponding to those ROIs, crop out ROIs (e.g., bounding boxes) corresponding to the ROIs, and send the cropped ROIs to the second camera via the network. The second camera is configured to receive the cropped ROIs from the first camera via the network, perform video analysis algorithms on the ROIs from the video frames captured by the first camera, generate video analysis results, and output these results via the network.
[0005] Another example can be seen in a method for load balancing video analytics processing among two or more networked cameras in a plurality of networked cameras, each networked camera including a camera for capturing a corresponding video stream and processing resources. This exemplary method includes a first camera among the plurality of networked cameras identifying an object of interest (ROI) in a video frame of a video stream captured by the first camera, and the first camera determining whether to perform a video analytics algorithm on the ROI to identify additional characteristics of the ROI. When it is determined that a video analytics algorithm should be performed on the ROI, the first camera determines whether it has sufficient idle processing resources to perform the video analytics algorithm on the ROI, and if so, the first camera performs the video analytics algorithm on the ROI. If the first camera does not have sufficient idle processing resources to perform the video analytics algorithm on the ROI, the first camera identifies a second camera among the plurality of networked cameras that has sufficient idle processing resources to perform the video analytics algorithm on the ROI, and sends a cropped region of interest (ROI) from a video frame of a video stream captured by the first camera to the second camera. Subsequently, the second camera receives the cropped region of interest (ROI) from the video frames of the video stream captured by the first camera, performs a video analysis algorithm on the cropped ROI of the video frames of the video stream captured by the first camera, thereby generating video analysis results, and returns the video analysis results to the first camera.
[0006] The foregoing description is provided to facilitate understanding of the innovative features unique to this disclosure and is not intended as a complete description. A full understanding of this disclosure can be obtained by considering the entire specification, claims, drawings, and abstract as a whole. Attached Figure Description
[0007] This disclosure will be more fully understood by taking into account the following description of various examples in conjunction with the accompanying drawings, in which:
[0008] Figure 1 This is a schematic block diagram illustrating an exemplary video surveillance system;
[0009] Figure 2A and Figure 2B This is a flowchart illustrating an exemplary method for load balancing video analytics processing across two or more cameras;
[0010] Figure 3 This is a flowchart illustrating an exemplary method for load balancing video analytics processing across two or more cameras; and
[0011] Figure 4 This is a flowchart illustrating an exemplary method.
[0012] While this disclosure is subject to various modifications and alternatives, its details have been shown by way of example in the accompanying drawings and will be described in detail. However, it should be understood that this disclosure is not intended to limit it to the specific examples described. Rather, it is intended to cover all modifications, equivalents, and alternatives that fall within the substance and scope of this disclosure. Detailed Implementation
[0013] The following description should be read with reference to the accompanying drawings, in which the same elements in the different drawings are numbered in the same manner. The drawings are not necessarily drawn to scale and depict examples that are not intended to limit the scope of this disclosure. Although examples of various elements are illustrated, those skilled in the art will recognize that many of the examples provided have suitable alternatives that can be utilized.
[0014] This document assumes that all numbers are modified by the term “about” unless otherwise explicitly stated. Expressions of numerical ranges using endpoints include all numbers contained within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, and 5).
[0015] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless otherwise expressly stated. As used in this specification and the appended claims, the term “or” is generally used in its meaning to include “and / or” unless otherwise expressly stated.
[0016] It should be noted that references to "one embodiment," "some embodiments," or "other embodiments" in the specification indicate that the described embodiments may include specific features, structures, or characteristics, but each embodiment need not necessarily include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in connection with an embodiment, it is conceivable that, whether explicitly described or not, that feature, structure, or characteristic may be applied to other embodiments, unless otherwise expressly stated otherwise.
[0017] Figure 1 This is a schematic block diagram illustrating an exemplary video surveillance system 10. The exemplary video surveillance system 10 includes a plurality of cameras 12, labeled 12a, 12b, and through 12n. The video surveillance system 10 may include any number of cameras 12, and in some cases may include dozens, hundreds, or even thousands of cameras 12. Each camera among the cameras 12 may be capable of capturing a video stream and performing one or more different video analysis algorithms on the captured video stream. Each camera among the cameras 12 may be considered to be configured to communicate via a network 14, and therefore may be considered a networked camera. The network 14 may be a LAN (Local Area Network) or a WAN (Wide Area Network). In some cases, the network 14 may include both a LAN and a WAN via one or more network components. In some cases, the network 14 may include the Internet. The exemplary video surveillance system 10 includes an NVR (Network Video Recorder) 16 connected via the network 14 and configured to store video streams provided to the NVR 16 by any of the cameras 12.
[0018] Camera 12a can be considered a first camera, and camera 12b can be considered a second camera. It should be understood that the names "first" and "second" are arbitrary and can refer to any two of the cameras in the 12 cameras. For example, camera 12a can be configured to capture a video stream and process the captured video stream. The captured video stream can be processed to identify regions of interest (ROIs) in the video frames corresponding to motion within the video frames, and in some cases, to identify objects of interest (ROIs) corresponding to the ROIs in the video frames. The captured video stream can be processed to crop out regions of interest (ROIs) corresponding to the ROIs in the video frames, and the cropped ROIs of the video frames can be sent to the second camera (such as camera 12b) via network 14. Camera 12b can then be configured to receive the cropped ROIs of the video frames from the first camera 12a via network 14, and perform a video analysis algorithm on the cropped ROIs of the video frames from the video stream captured by the first camera 12a, thereby producing a video analysis result. Camera 12b can then output the video analysis result via network 14. In some cases, camera 12b may be configured to output video analysis results to camera 12a via network 14. In some cases, camera 12b may be configured to output video analysis results to NVR 16 via network 14. In some cases, camera 12b may be configured to output video analysis results to camera 12a via network 14 and also to NVR 16 via network 14.
[0019] In some cases, video analysis results may include one or more tags describing the object of interest. These tags may be associated with timestamps in the video stream. A second camera 12b may be configured to output the video analysis results to camera 12a via network 14, and camera 12a may be configured to integrate one or more tags into the live stream of the video stream captured by camera 12a. This may include overlaying the tags onto the live stream adjacent to the object of interest and aligning them substantially temporally based on timestamps. In some cases, there may be a delay of one or two frames before the tags appear in the live stream, but this is acceptable in many cases.
[0020] In some cases, camera 12b may be configured to output video analysis results to NVR 16 via network 14, which records the video stream captured by camera 12a. NVR 16 may be configured to receive the video analysis results from camera 12b (or camera 12a) and integrate the one or more tags into the recorded video stream captured by camera 12a. This may include overlaying tags onto video frames of the recorded video stream adjacent to the object of interest and aligning them temporally based on timestamps. In this case, there may be no delay in the tags appearing in the recorded video stream.
[0021] In some cases, camera 12a may be configured to send metadata to camera 12b, which identifies from a plurality of predetermined video analysis algorithms a video analysis algorithm to be performed by the second camera 12b on cropped regions of interest (ROIs) of video frames captured by camera 12a. In some cases, camera 12a may include processing resources with a current resource utilization level, and camera 12a may be configured to determine, before sending the cropped regions of interest (ROIs) of video frames from the video stream to camera 12b, that the current resource utilization level of camera 12a's processing resources exceeds a threshold utilization level. In some cases, the threshold utilization level may depend on the video analysis algorithm to be performed on the cropped regions of interest (ROIs) of video frames from the video stream. For example, an object classification video analysis algorithm may require more processing resources than a face recognition video analysis algorithm.
[0022] Figure 2A and Figure 2B This is a flowchart illustrating an exemplary method 18 for load balancing video analytics processing among two or more networked cameras (such as cameras 12a and 12b) in a plurality of networked cameras, each networked camera including a camera for capturing a corresponding video stream. Exemplary method 18 includes a first camera among the plurality of networked cameras identifying objects of interest in video frames of the video stream captured by that first camera, as indicated in box 20. In some cases, identifying objects of interest in video frames of the video stream captured by the first camera may include identifying pixel regions in the video frames of the video stream that differ from corresponding pixels in a reference video frame, and identifying objects of interest in video frames as pixel regions in the video frames that differ from corresponding pixels in a reference video frame.
[0023] The first camera crops out the Region of Interest (ROI) corresponding to the object of interest from the video frames of the video stream, where the cropped region includes less than the entire video frames of the video stream, as indicated in box 22. The first camera sends the cropped ROI of the video frames of the video stream to a second camera among the networked cameras, as indicated in box 24. The second camera then performs a video analysis algorithm on the cropped ROI of the video frames of the video stream, thereby producing a video analysis result, as indicated in box 26. The second camera then sends the video analysis result to the first camera among the networked cameras, as indicated in box 28.
[0024] In some cases, a first camera among multiple networked cameras may include processing resources with a current resource utilization level, and the first camera may determine that the current resource utilization level of its processing resources exceeds a threshold utilization level before sending cropped regions of interest (ROIs) of video frames from the video stream to a second camera among the multiple networked cameras. In some cases, the threshold utilization level may depend on the video analysis algorithm to be performed on the cropped ROIs of the video frames from the video stream. For example, an object classification video analysis algorithm may require more processing resources than a face recognition video analysis algorithm. In some cases, when the first camera determines that the current resource utilization level of its processing resources does not exceed the threshold utilization level, the first camera performs a video analysis algorithm on the cropped ROIs of the video frames from the video stream and does not send the cropped ROIs of the video frames from the video stream to a second camera among the multiple networked cameras.
[0025] In some cases, and continue Figure 2B Example method 18 may include a first camera among a plurality of networked cameras classifying an object of interest into one of a plurality of categories, as indicated in box 30. For example, the first camera may perform a coarse object classification of the object of interest (e.g., a person, a car, a dog, etc.). The first camera among the plurality of networked cameras may send the classification of the object of interest, along with the cropped region of interest (ROI) of the video frames of the video stream, to a second camera, as indicated in box 32. In some cases, the first camera among the plurality of networked cameras may send metadata to the second camera that identifies from a plurality of predetermined video analysis algorithms the video analysis algorithm to be performed by the second camera on the cropped region of interest (ROI) of the video frames of the video stream, as indicated in box 34.
[0026] In some cases, each of the multiple networked cameras may include processing resources with a corresponding current resource utilization level, and each of the multiple networked cameras may make its corresponding current resource utilization level known to the other networked cameras. In some cases, a first camera among the multiple networked cameras may select a second camera from the multiple networked cameras based at least in part on the current resource utilization level of the second camera, as indicated in box 38. In some cases, the first camera among the multiple networked cameras may convert the cropped region of interest (ROI) of the video frame to grayscale before sending it to the second camera among the multiple networked cameras. This can reduce the network bandwidth required to send the cropped ROI to the second camera over the network and can reduce the processing resources of the second camera required to perform a specified video analysis algorithm. For many video analysis algorithms, there is almost no loss of accuracy when using grayscale video frames compared to panchromatic video frames.
[0027] In some cases, video analysis results sent from a second camera to a first camera may include one or more tags describing one or more characteristics of an object of interest, wherein the first camera integrates these tags into a live stream of video captured by the first camera. In some cases, the video analysis results sent from the second camera may include one or more tags describing one or more characteristics of an object of interest, and the first camera may be operatively coupled to a network video recorder (NVR) that records the video stream captured by the first camera. The NVR can receive the video analysis results and can integrate one or more tags into the recorded video stream captured by the first camera.
[0028] Figure 3This is a flowchart illustrating an exemplary method 40 for load balancing video analytics processing among two or more networked cameras in a plurality of networked cameras, each networked camera including a camera for capturing a corresponding video stream and processing resources. Exemplary method 40 includes a first camera among the plurality of networked cameras identifying objects of interest (POIs) in video frames of the video stream captured by the first camera, as indicated in box 42. In some cases, the first camera may perform a coarse object classification of the POIs (e.g., people, cars, dogs, etc.). The first camera determines whether to perform a video analytics algorithm on the POIs to identify additional characteristics of the POIs, as indicated in box 44. For example, the first camera may determine that when the POIs are classified as people, a facial recognition video analytics algorithm should be performed on the POIs to identify the person. When it is determined that a video analytics algorithm should be performed on the POIs, the first camera determines whether it has sufficient idle processing resources to perform the video analytics algorithm on the POIs, as indicated in box 46. If so, the first camera performs the video analytics algorithm on the POIs, as indicated in box 48.
[0029] However, if it is determined that the first camera does not have sufficient idle processing resources, the first camera identifies a second camera among multiple networked cameras that has sufficient idle processing resources to perform a video analysis algorithm on the object of interest, and sends the cropped region of interest (ROI) from the video frames of the video stream captured by the first camera to the second camera, as indicated in box 50. The second camera then takes several actions, as indicated in box 52. The second camera receives the cropped ROI from the video frames of the first camera, as indicated in box 52a. The second camera then performs a video analysis algorithm on the cropped ROI from the video frames of the video stream captured by the first camera, thereby producing a video analysis result, as indicated in box 52b. In this example, the second camera returns the video analysis result to the first camera, as indicated in box 52c.
[0030] Figure 4This is a flowchart illustrating exemplary method 54. It should be understood that multiple cameras 12 may exist, here labeled 12a, 12b, 12c, 12d, and so on, up to 12n. Each of the cameras 12 includes a database 56 that tracks the available or idle CPU (Central Processing Unit) and / or GPU (Graphics Processing Unit) processing resources for each of the multiple cameras 12. This information can be repeatedly updated via a network connecting the multiple cameras 12. Method 54 begins at start box 58. As indicated at box 60, the algorithm for camera 12a is enabled. At box 62, it is calculated how many bounding boxes have been created (e.g., how many objects of interest have been identified), and the processing resources required to process each bounding box, as well as the cumulative processing resources required to process all N bounding boxes, are calculated. At decision box 64, it is determined whether the current camera 12a has sufficient idle processing resources. If so, the current camera 12a processes the data itself.
[0031] However, if the current camera 12a does not have sufficient idle processing resources, control proceeds to decision box 68, where it is determined which of the "N" cameras has sufficient processing resources to process one or more bounding boxes within the bounding box. Control then proceeds to box 70, where each bounding box is sent to the corresponding idle camera for processing, and database 56 is updated. If no idle camera can be found for some bounding boxes (temporarily isolated bounding boxes), control proceeds to box 72, where temporarily isolated bounding boxes are sequentially sent to the cameras as processing resources become available. If a particular camera has more idle processing resources, more than one bounding box can be sent to that camera. If, for some reason, not all bounding boxes can be sent, an update to database 56 is indicated at box 74, and control returns to the start box 58.
[0032] Although several exemplary embodiments of this disclosure have been described thus, those skilled in the art will readily understand that other embodiments can be made and used within the scope of the appended claims. However, it should be understood that this disclosure is merely illustrative in many respects. Changes may be made to details, particularly those relating to shape, size, arrangement of parts, and exclusion and order of steps, without departing from the scope of this disclosure. The scope of this disclosure is, of course, defined by the language of the appended claims.
Claims
1. A method for load balancing video analytics processing among two or more networked cameras (12a, 12b, 12c) in a plurality of networked cameras, each networked camera including a camera for capturing a corresponding video stream, the method comprising: The first camera among the plurality of networked cameras identifies the object of interest in a video frame of the video stream captured by the first camera; The region of interest (ROI) corresponding to the object of interest is cropped from the video frames of the video stream, wherein the cropped region includes less than all of the video frames of the video stream; The region of interest (ROI) of the video frame of the video stream is sent to the second camera among the plurality of networked cameras; The second camera among the plurality of networked cameras performs a video analysis algorithm on the cropped region of interest (ROI) of the video frame of the video stream, thereby generating a video analysis result; as well as The second camera among the plurality of networked cameras sends the video analysis results to the first camera among the plurality of networked cameras.
2. The method of claim 1, wherein identifying the object of interest in the video frames of the video stream captured by the first camera comprises: Identify pixel regions in the video frames of the video stream that are different from corresponding pixels in the reference video frame; as well as The object of interest in the video frame is identified as corresponding to a pixel region in the video frame of the video stream that is different from the corresponding pixel in the reference video frame.
3. The method according to claim 1, wherein the method comprises: The first camera among the plurality of networked cameras classifies the object of interest into one of the plurality of categories; as well as The first camera among the plurality of networked cameras sends the classification of the object of interest, along with the cropped region of interest (ROI) of the video frame of the video stream, to the second camera.
4. The method according to claim 1, wherein the method comprises: The first camera among the plurality of networked cameras sends metadata to the second camera, the metadata identifying from a plurality of predetermined video analysis algorithms the video analysis algorithm to be executed by the second camera on the cropped region of interest (ROI) of the video frame of the video stream.
5. The method of claim 1, wherein the first camera among the plurality of networked cameras includes processing resources having a current resource utilization level, and wherein the first camera determines that the current resource utilization level of the processing resources of the first camera exceeds a threshold utilization level before sending the cropped region of interest (ROI) of the video frame of the video stream to the second camera among the plurality of networked cameras.
6. The method according to any one of claims 1 to 5, wherein each of the plurality of networked cameras includes processing resources having a corresponding current resource utilization level, and wherein each of the plurality of networked cameras makes its corresponding current resource utilization level known to all other networked cameras in the plurality of networked cameras.
7. The method according to any one of claims 1 to 5, wherein the method comprises: Before sending the cropped region of interest (ROI) of the video frame of the video stream to the second camera among the plurality of networked cameras, the first camera among the plurality of networked cameras converts the cropped region of interest (ROI) of the video frame to grayscale.
8. The method according to any one of claims 1 to 5, wherein the video analysis result sent from the second camera to the first camera includes one or more tags describing one or more characteristics of the object of interest, wherein the first camera integrates the one or more tags into a live stream of the video stream captured by the first camera.
9. A monitoring system, the monitoring system comprising: First camera (12a); Network (14); A second camera (12b) is operatively coupled to the first camera via the network; The first camera is configured as follows: Capture video stream; Process the video stream to: Identify motion regions in video frames of the video stream that correspond to motion in the video frames; Identify the object of interest corresponding to the motion region in the video frame; Cropping out the region of interest (ROI) in the video frame that corresponds to the object of interest; The cropped region of interest (ROI) of the video frame is sent to the second camera via the network; The second camera is configured as follows: The cropped region of interest (ROI) of the video frame received from the first camera is received via the network. A video analysis algorithm is performed on the cropped region of interest (ROI) of the video frame of the video stream captured by the first camera to produce video analysis results; as well as The video analysis results are output via the network.
10. A method for load balancing video analytics processing among two or more networked cameras (12a, 12b, 12c) in a plurality of networked cameras, each networked camera including a camera for capturing a corresponding video stream and processing resources, the method comprising: The first camera among the plurality of networked cameras identifies the object of interest in a video frame of the video stream captured by the first camera; The first camera determines whether to perform a video analysis algorithm on the object of interest to identify additional characteristics of the object of interest; When it is determined that the video analysis algorithm will be executed on the object of interest, the first camera determines whether it has sufficient idle processing resources to execute the video analysis algorithm on the object of interest. If yes, the first camera executes the video analysis algorithm on the object of interest; otherwise: The first camera identifies the second camera among the plurality of networked cameras that has sufficient idle processing resources to perform the video analysis algorithm on the object of interest, and sends the cropped region of interest (ROI) from the video frames of the video stream captured by the first camera to the second camera; Second camera: Receive the cropped region of interest (ROI) from the video frame of the first camera; The video analysis algorithm is performed on the cropped region of interest (ROI) of the video frame of the video stream captured by the first camera to generate video analysis results; as well as The video analysis results are returned to the first camera.