Object Detection
By evaluating the correlation between the image and the previous frame in the video frame and comparing the energy consumption, deciding whether to perform object detection or tracking, the problem of inefficient video object detection and tracking in the prior art is solved, and efficient and low-energy real-time detection and tracking are achieved.
Patent Information
- Application Number
- CN201980101126.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-10-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-10-02
AI Technical Summary
Existing video object detection and tracking methods are inefficient in complex scenarios, especially when real-time high-precision detection is required, the computational complexity and energy consumption are high.
By determining the correlation between the image and the previous image in a video frame and deciding whether to perform an object detection or tracking process based on the comparison of energy consumption, the algorithm is optimized to reduce unnecessary calculations.
It realizes the ability to maintain detection accuracy while reducing energy consumption and improving efficiency, and can perform object detection and tracking tasks in real time.
Smart Images

Figure CN114503168B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to object detection, and in particular to the detection of objects in video images. Background Art
[0002] It is well known to use a convolutional neural network (CNN) to detect objects in an image, but this is a computationally intensive task. Each complete pass of a detection or inference algorithm may require billions of operations. Although one-stage object detection methods developed in recent years have shown promising accuracy and significantly reduced complexity compared to two-stage methods, the complexity is still high, and real-time operations with high accuracy can only be achieved using specialized graphics processing units (GPUs).
[0003] In contrast, tracking previously detected objects from one video frame to another may be a relatively low-complexity operation. Existing single-object trackers run at hundreds of frames per second (fps) using a general central processing unit and have convincing tracking results. However, tracking many objects at once may actually be less efficient than performing object detection.
[0004] The document "Detect or Track: Towards Cost-Effective Video Object Detection / Tracking", Luo, Hao, Wenxuan Xie, Xinggang Wang, and Wenjun Zeng, ArXiv:1811.05340 [cs.CV], November 13, 2018 discloses a system in which a scheduler makes a decision to use detection or tracking for each input frame and processes the frames accordingly. In the prior art system, a Siamese network is used to track multiple objects, and a decision to use tracking or detection for a new frame is made based on the output of a scheduler network, which in turn is based on the output of the Siamese network. The scheduler is trained using reinforcement learning. Summary of the Invention
[0005] According to one aspect of the present invention, there is provided a method for detecting an object in a video, the method comprising, for each image in the video:
[0006] determining a degree of correlation between the image and at least one previous image in an image sequence;
[0007] Obtain a first measure of the energy consumption required to perform an object detection process on an image, obtain a second measure of the energy consumption required to perform an object tracking process for tracking at least one object from a previous image, and compare the first measure of energy and the second measure; and
[0008] If the correlation is lower than a threshold level or the first measure of energy is lower than the second measure of energy, perform an object detection process on the image, otherwise perform an object tracking process on the image.
[0009] The method may include performing an object detection process using a neural network.
[0010] The method may include performing an object detection process using a one-stage object detector, or may include performing an object detection process using a multi-stage object detector.
[0011] The method may include obtaining a first measure of energy consumption based on a predetermined power consumption of a neural network.
[0012] The method may include using a single object tracker to perform object tracking.
[0013] The method may include obtaining a first measure of energy consumption based on information about the device to be used for performing the object detection process.
[0014] The method may include obtaining a second measure of energy consumption based on information about the device to be used for tracking.
[0015] The method may include obtaining a second measure of energy consumption based on the number of objects being tracked and their sizes.
[0016] The method may include using a regression function obtained from previous measurements of energy consumption to obtain a second measure of energy consumption.
[0017] The method may include performing an object detection process or an object tracking process in the compressed domain.
[0018] The method may include performing data association to calculate tracklets for each detected object.
[0019] The method may include determining the correlation between an image and at least one previous image by the following steps:
[0020] Determine a first number of matching features in the previous image;
[0021] Determine a second number of matching features in the current image; and
[0022] Determine the correlation based on the difference between the first number of matching features and the second number of matching features.
[0023] The first quantity of features can be an average of the respective quantities of features in each of the plurality of previous images.
[0024] The method can include determining a degree of correlation between the image and at least one previous image based on features extracted from the image.
[0025] The method can include extracting features by:
[0026] extracting features from the full image; or
[0027] extracting features from a downsampled image; or
[0028] dividing the image into blocks and extracting corresponding features from each block of the image.
[0029] The method can further include determining the degree of correlation between the image and at least one previous image by using additional sensor outputs and / or by using features from a compressed video.
[0030] According to another aspect, there is provided a computer program product including code for causing a programmed processor to execute the method according to the first aspect.
[0031] According to another aspect, there is provided a video analysis processing device including:
[0032] an input for receiving video; and
[0033] a processor configured to execute a method that, for each image in the video, includes:
[0034] determining a degree of correlation between the image and at least one previous image in an image sequence;
[0035] obtaining a first measure of the energy consumption required to perform an object detection process on the image, obtaining a second measure of the energy consumption required to perform an object tracking process for tracking at least one object from a previous image, and comparing the first measure of energy and the second measure of energy; and
[0036] if the degree of correlation is below a threshold level or the first measure of energy is lower than the second measure of energy, performing an object detection process on the image, otherwise performing an object tracking process on the image.
[0037] According to another aspect, there is provided a video analysis device including the image analysis processing device according to the foregoing aspect and further including a camera.
[0038] The disclosed embodiments can improve the energy consumption of object detection by determining whether to use object detection or object tracking at each stage and considering energy consumption when making the determination. From the perspective of energy consumption, the process can be efficient while maintaining detection accuracy and fast enough to be executed in real time. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] To better understand the present invention and illustrate how it is implemented, reference will now be made, by way of example, to the accompanying drawings, in which:
[0040] Figure 1 is a block diagram showing the main functional components of a device configured for video analysis.
[0041] Figure 2 schematically shows the operations of a method for detecting an object.
[0042] Figure 3 is a flowchart showing an embodiment of a method according to the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] Figure 1 is a block diagram showing the main functional components of a video analysis device 10.
[0044] In the present embodiment, the video analysis device 10 includes a camera sensor 20 as a source of video images. However, the video analysis device 10 can operate on video images received from outside the video analysis device.
[0045] The video image received from the source is passed to the input 28 of a video analysis processing device 30, which exists in the form of a data processing and control unit and includes a processor 32 and a memory 34.
[0046] The processor 32 controls the operation of the video analysis processing device 30 and / or the video analysis device 10. The processor 32 can include one or more microprocessors, hardware, firmware, or a combination thereof. In particular, as described in more detail below, the processor 32 can be configured to have neural network functionality.
[0047] The memory 34 can include volatile and non-volatile memories for storing computer program code and data required for the operation by the processor 32.
[0048] The memory 34 may include any tangible, non-transitory computer-readable storage medium for storing data, including electronic, magnetic, optical, electromagnetic, or semiconductor data storage. The memory 34 stores at least one computer program including executable instructions that configure the processor 32 to implement the methods and processes described herein. Generally, computer program instructions and configuration information are stored in non-volatile memory (such as read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory). Temporary data generated during operation may be stored in volatile memory (such as random access memory (RAM)). In some embodiments, the computer program for configuring the processor 32 as described herein may be stored on a removable memory (such as a portable compact disc, portable digital video disc, or other removable medium). The computer program may also be implemented in a carrier such as an electronic signal, optical signal, radio signal, or computer-readable storage medium).
[0049] The video analysis device 10 with additional functions can be incorporated into any type of device, which will not be elaborated here. Generally, the method is particularly useful in devices with limited computing power and / or power constraints. For example, the video analysis device can be incorporated into a smartphone, in which case it will include the conventional features of the smartphone, such as user interface features such as a touch screen, microphone, and speaker, as well as communication features such as an antenna, RF transceiver circuit, etc. The method is particularly useful when the device is operating in a low-power mode. As another example, the video analysis device 10 can be incorporated into an Internet of Things (IoT) device and operate with strict power limitations.
[0050] Figure 2 The operations of a method for detecting an object in the video analysis device 10 are schematically illustrated.
[0051] Specifically, as shown in block 50, the video analysis device 10 generates information related to the correlation between the current image I t and at least one previous image I t-1 in the image sequence. In block 52, the video analysis device 10 obtains a first measure of the energy consumption required to perform an object detection process on the image and obtains a second measure of the energy consumption required to perform an object tracking process for tracking at least one object from the previous image, where the latter is based on the number N t-1 of objects and the size of the region to be tracked in the previous frame.
[0052] In block 54, the video analysis device compares the first measure of energy and the second measure of energy and determines that if the correlation is below a threshold level or the first measure of energy is lower than the second measure of energy, then an object detection process 56 should be performed on the image, or otherwise an object tracking process 58 should be performed on the image.
[0053] Figure 3 It is a more detailed flowchart showing a method for performing object detection.
[0054] When the system starts, the method begins at step 70.
[0055] At step 72, an image sequence is obtained. In the device 10 as Figure 1 shown, the image sequence includes a sequence of video frames obtained from the camera sensor 20.
[0056] An object detection process is performed on the first frame in the sequence, as described in more detail below. Then, the steps of the method are applied to each successive image or frame in the sequence.
[0057] At step 74, image processing may be performed.
[0058] In one embodiment, selected image processing steps from an image signal processor (ISP) are applied to the image. For example, image processing steps such as demosaicking the raw image may be performed, and / or contrast enhancement may be applied.
[0059] At step 76, image feature points are extracted from the image. For example, each image may first be downsampled (e.g., to a resolution of 300*300 pixels). Features may be extracted by using an Oriented FAST and Rotated BRIEF (ORB) feature detector. Other feature extraction algorithms may also be used for feature extraction, such as Scale-Invariant Feature Transform (SIFT) or Speeded-Up Robust Features (SURF) feature detectors. If the received video is in a compressed format, it may be necessary to partially decompress it to allow feature extraction.
[0060] At step 78, the degree of correlation between the image and at least one previous image in the image sequence is determined.
[0061] For example, frame correlation analysis may include analyzing the correlation between frames based on the deviation of the number of matching feature points between the current frame and the previous frame.
[0062] The number of matching points is obtained by matching the current frame with the previous frame in the feature space.
[0063] For example, if the number of matching feature points of the i-th video frame is denoted as N i , then the moving average of N i is calculated as N avg . In addition, the deviation value d i for each frame is calculated as d i = abs(N i - N avg) The moving average d of the deviation avg is calculated as an exponential average, e.g., d avg = 0.9 * d avg + 0.1 * d i .
[0064] Then, if d i > ρd avg , where the parameter ρ is a factor that adjusts the sensitivity of the calculation and is assigned a value of 2.7 in one embodiment, this is regarded as an indication that there has been a significant change in the image sequence (e.g., a change in the scene and the need to perform a new object detection process).
[0065] In the above embodiment, the correlation analysis operates by considering the correlation of the entire image (i.e., the number of matching points). In an alternative embodiment, this global point matching is replaced by a process of matching image segments (such as blocks).
[0066] Accordingly, the image is divided into multiple segments, and in step 76, the number of matching feature points is calculated for each segment and the number of matching feature points for the entire image is counted. Then, in step 78, a similar calculation to the above can be performed.
[0067] Such a calculation with large blocks can accelerate the calculation. The size of the blocks needs to be carefully selected because if there are sufficiently large blocks and smaller objects in the scene, then extracting and matching key points may return incorrect associations.
[0068] The correlation analysis can also include determining the degree of correlation by using the output from additional sensors (such as depth cameras, dynamic vision sensors, inertial sensors, etc.) and by using the encoded features (e.g., motion vectors) from compressed video. This additional information can be used to more precisely determine the degree of correlation. For example, if the inertial measurement unit on the device detects a sudden movement, this can be used to determine that a low correlation between images can be expected.
[0069] The correlation analysis algorithm should be lightweight and its implementation should be efficient, i.e., the complexity of the correlation analysis should be much lower than the complexity of object detection to minimize the energy usage of the overhead.
[0070] In step 80, a first measure or estimate of the energy consumption required to perform an object detection process on the image is obtained, and a second measure or estimate of the energy consumption required to perform an object tracking process to track at least one object from a previous image is obtained. These energy consumption measures or estimates are obtained from the information about the previous image.
[0071] For example, a first measure of energy consumption can be obtained based on information about the hardware device to be used for performing the object detection process. This information can be obtained in advance by measuring the energy consumption of the device during a trial detection process. Additionally or alternatively, a first measure of energy consumption can be obtained based on information about the computational complexity of the object detection process to be used. For example, if a neural network is to be used to perform the object detection process, a first measure of energy consumption can be obtained based on the predetermined power consumption of the neural network.
[0072] Assuming objects are tracked individually, a second measure of energy consumption can be calculated as the sum of the energies of all the tracked objects. In another embodiment, the tracker used is a Kernelized Correlation Filter (KCF), and the energy consumption of each tracked object is related to the size m*n of the tracking box and can be fitted to a polynomial / regression function. Additionally or alternatively, a second measure of energy consumption can be obtained based on information about the hardware device to be used for performing the object tracking process.
[0073] Thus, for example, a particular hardware device can be used to track pixels on multiple occasions. For example, it can be assumed through a polynomial function that the energy consumption (y) is related to the number of pixels tracked (x), such as:
[0074] y = bx 2 + cx + d,
[0075] Then the least squares method can be used to fit this function to the values obtained on these occasions to obtain suitable values for b, c, and d.
[0076] Then, an estimate of the energy consumption can be obtained on future occasions when the number of pixels to be tracked has been determined.
[0077] In one embodiment, estimates of energy consumption for both tracking and detection can be generated by measuring the energy usage on the device over multiple tracking and detection iterations. Based on the measured energy usage and the information about the correlation between frames determined in step 78, the regression function as described above can be obtained for the actual energy consumption on the device and used to derive subsequent energy consumption estimates.
[0078] The process then proceeds to step 82a, in which a first part of the determination of whether to perform an object detection process or an object tracking process on the current image is made.
[0079] If the degree of correlation obtained in step 78 is lower than a threshold, which is selected based on the degree of correlation required to be confident that the object tracking process will succeed, it is determined that the object detection process should be performed, and the method proceeds to step 84.
[0080] If the relevance obtained in step 78 is higher than the threshold, i.e., it is determined that the object tracking process will be able to detect the object present in the image, the process proceeds to step 82b, where a first measure or estimate of the energy consumption required to perform the object detection process on the image is compared with a second measure or estimate of the energy consumption required to perform the object tracking process. If it is determined that the estimated energy consumption of the detection process is less than the estimated energy consumption of the tracking process, it is determined that the object detection process should be performed and the method proceeds to step 84, even though the relevance obtained in step 78 is higher than the threshold, which means that the object tracking process will be successful, i.e., the object tracking process will be able to detect the object present in the image.
[0081] If the relevance obtained in step 78 is higher than the threshold and it is determined that the estimated energy consumption of the detection process is greater than the estimated energy consumption of the tracking process, it is determined that the object tracking process should be performed and the method proceeds to step 86.
[0082] In the illustrated embodiment, energy estimates are made in all cases, but they are compared in step 82b only when the relevance exceeds the threshold. In other embodiments, energy estimates are made and then compared only when the relevance exceeds the threshold. In other embodiments, energy estimates are made and compared in all cases, but the result of the comparison is useful only when the relevance exceeds the threshold.
[0083] As described above, in step 84, the object detection process is performed.
[0084] The object detection process can use a neural network. The object detection process can be a one-stage object detector, such as a Single Shot Detector (SSD), e.g., MobileNet-SSD, or a Deeply Supervised Object Detector (DSOD), e.g., Tiny-DSOD. Alternatively, the object detection process can be a two-stage object detector or a multi-stage object detector.
[0085] To perform object detection on an image in the compressed domain, a partial decompression of the compressed video bitstream is performed, and prediction mode, block segmentation information, DCT coefficients, etc. are used as the input format for the neural network for object detection.
[0086] As described above, in step 86, the object tracking process is performed.
[0087] In some embodiments, the object tracking process uses a single object tracker, although in other embodiments a multi-object tracker may be used. A low-complexity object tracking algorithm may be used. For example, a suitable object tracking algorithm uses a kernel correlation filter, as described in "High-Speed Tracking with Kernelized Correlation Filters", Henriques, Joao F., Rui Caseiro, Pedro Martins, and Jorge Batista, IEEE Transactions on Pattern Analysis and Machine Intelligence 37, no. 3 (March 1, 2015): 583–596.
[0088] In some embodiments, as described above, the features extracted in step 76 of the method may be reused for object tracking.
[0089] In one embodiment, object tracking is performed in the compressed video domain.
[0090] To perform object tracking on an image in the compressed domain, a partial decompression of the compressed video bitstream is performed, and the object tracking is based on motion vectors extracted from the video codec by using a method such as a spatio-temporal Markov random field.
[0091] After the object detection process in step 84 or the object tracking process in step 86 has been performed, the method may proceed to step 88, where data association of the object detection and tracking results may be performed. This step is typically skipped only for the object detection operation. However, this module may be enabled to provide the system with the possibility of computing a tracklet (i.e., a segment of the trajectory followed by a moving object) for each object.
[0092] The data association process may use the Hungarian algorithm to create the trajectories of the tracked objects.
[0093] The above method is then executed until all the images of the video have been processed, and then the method proceeds to step 90, where it terminates.
[0094] In this way, a method for performing object detection is described, which can be executed in an efficient manner by determining when to track and when to detect in a sequence of video frames, and thus the energy consumption of the device can be reduced.
[0095] It should be noted that the above embodiments illustrate rather than limit the present invention, and those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. The word "comprising" does not exclude the presence of elements or steps other than those listed in the claims, "a" or "an" does not exclude a plurality, and a single feature or other unit may perform the functions of a plurality of units described in the claims. Any reference signs in the claims should not be construed as limiting their scope.
Claims
1. A method for detecting an object in a video, the method comprising, for each image in the video: Determine (78) a degree of correlation between the image and at least one previous image in an image sequence, wherein determining the degree of correlation is based on a difference between a number of matching features of the image and a number of matching features of the at least one previous image; Obtain (80) a first measure of energy consumption required to perform an object detection process on the image, obtain a second measure of energy consumption required to perform an object tracking process for tracking at least one object from a previous image, and compare (82b) the first measure and the second measure of energy consumption; If the degree of correlation is below a threshold level or the first measure of energy consumption is lower than the second measure of energy consumption, then perform the object detection process (84) on the image; And If the degree of correlation is above the threshold level and the first measure of energy consumption is greater than the second measure of energy consumption, then perform the object tracking process (86) on the image.
2. The method according to claim 1, comprising performing the object detection process using a neural network.
3. The method according to claim 2, comprising performing the object detection process using a one-stage object detector.
4. The method according to claim 2, comprising performing the object detection process using a multi-stage object detector.
5. The method according to claim 2, Comprising: Obtain the first measure of energy consumption based on a predetermined power consumption of the neural network.
6. The method according to claim 1, comprising using a single object tracker to perform the object tracking.
7. The method according to claim 1, comprising obtaining the first measure of energy consumption based on information about a device to be used for performing the object detection process.
8. The method according to claim 1, comprising obtaining the second measure of energy consumption based on information about a device to be used for tracking.
9. The method according to claim 1, comprising obtaining the second measure of energy consumption based on the number and size of the objects being tracked.
10. The method according to claim 9, comprising using a regression function obtained from previous energy consumption measurements to obtain the second measure of energy consumption.
11. The method according to claim 1, Characterized in that Comprising performing the object detection process or the object tracking process in a compression domain.
12. The method according to claim 1, comprising performing data association (88) to calculate a tracklet for each detected object.
13. The method according to claim 1, comprising determining the degree of correlation between the image and at least one previous image by the following steps: Determine a first number of matching features in the previous image; Determine a second number of matching features in the current image; and Determine the degree of correlation based on a difference between the first number of matching features and the second number of matching features.
14. The method according to claim 13, wherein the first number of features is an average of the respective numbers of features in each of a plurality of previous images.
15. The method according to claim 1, comprising determining the relevance between the image and the at least one previous image based on features extracted from the image.
16. The method according to claim 15, comprising extracting the features by the steps of: Extracting features from the complete image.
17. The method according to claim 15, comprising extracting the features by the steps of: Extracting features from the downsampled image.
18. The method according to claim 15, comprising extracting the features by the steps of: Dividing the image into blocks and extracting corresponding features from each block of the image.
19. The method according to any one of the preceding claims, further comprising determining the relevance between the image and the at least one previous image by using additional sensor outputs and / or by using features from a compressed video.
20. A computer program product comprising code for causing a programmed processor to execute the method according to any one of claims 1 to 19.
21. A video analysis processing device (30), comprising: An input (28) for receiving video; and A processor (32) configured to execute a method which, for each image in the video, comprises: Determining the relevance between the image and at least one previous image in an image sequence, wherein determining the relevance is based on the difference between the number of matching features of the image and the number of matching features of the at least one previous image; Obtaining a first measure of the energy consumption required to perform an object detection process on the image, obtaining a second measure of the energy consumption required to perform an object tracking process for tracking at least one object from a previous image, and comparing the first measure and the second measure of energy; and If the relevance is below a threshold level or the first measure of energy consumption is lower than the second measure of energy consumption, performing the object detection process on the image; If the relevance is above the threshold level and the first measure of energy consumption is greater than the second measure of energy consumption, performing the object tracking process on the image.
22. A video analysis device (10) comprising the image analysis processing means (30) according to claim 21, and further comprising a camera (20).
Citation Information
Patent Citations
Target tracking method and device
CN106803263A
Multi-tracker object tracking
US20150055821A1
System and method of hybrid tracking for match moving
US20180137632A1