Methods, devices, and systems for detecting and tracking objects in captured video using convolutional neural networks
The detection and tracking model built using convolutional neural networks solves the problem of missed lesion diagnosis in endoscopic examinations, and achieves high-precision real-time detection and tracking of early gastric cancer and colorectal cancer, thereby reducing the missed diagnosis rate.
Patent Information
- Application Number
- CN202280002183.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-06-07
- Filing Date
- 2022-07-13
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2042-07-13
AI Technical Summary
In existing endoscopic examinations, the lesion detector has a high rate of missed diagnosis due to artifacts when processing endoscopic videos, especially for early gastric cancer and colorectal cancer. Current technology is not able to effectively detect and track lesions in real time.
A convolutional neural network is used to construct a detection and tracking model. Image data is generated by the processor of a video surveillance device. The detection and tracking models are used to perform target detection and tracking. By combining detection scoring enhancement and tracking reliability estimation, the detection accuracy and tracking stability are improved, and the impact of artifacts on detection is reduced.
It improves the detection accuracy and tracking stability of lesions during endoscopic examinations, reduces the rate of missed diagnoses, and enables effective real-time detection and tracking of early gastric and colorectal cancers.
Smart Images

Figure CN115335860B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to target detection and tracking techniques in captured videos. More specifically, this invention relates to methods, apparatus, and systems for detecting and tracking objects in captured videos using convolutional neural networks. Background Technology
[0002] Gastrointestinal endoscopy is a commonly used method for differentiating between gastric and colorectal cancer. Early endoscopic examination of gastric and colorectal cancer is the most effective way to reduce cancer mortality. In a study [S. Menon and N. Trudgill, “How commonly is upper gastrointestinal cancer missed at endoscopy? A meta-analysis,” Endosc Int Open, vol. 2, no. 2, pp. E46-E50, 2014], a meta-analysis of 3,787 patients with upper gastrointestinal (UGI) cancer showed that 11.3% of UGI cancers had been missed by endoscopy within 3 years prior to diagnosis. Furthermore, the rate of polyps missed by colonoscopy is approximately 20% [van Rijn, JC et al. Polyp miss rate determined by tandem colonoscopy: a systematic review. Am. J. Gastroenterol. 101, 343–350, 2006].
[0003] Machine learning-based lesion detectors are typically trained on qualified images and can process and interpret qualified endoscopic images very effectively. However, directly applying such lesion detectors to endoscopic videos may produce unreliable results because artifacts are very common in endoscopic videos, such as underexposure / overexposure, motion blur, video defocus, fluids, bubbles, specular reflections, and floating objects.
[0004] Developing an artificial intelligence algorithm to detect early-stage gastric and colorectal cancer and help prevent lesions from being missed during endoscopy, especially by detecting and tracking lesions (target objects) in real time during video endoscopy, is crucial. Summary of the Invention
[0005] According to one aspect of the present invention, a computer-implemented method for detecting and tracking target objects in captured video using a convolutional neural network via a video surveillance device includes: a processor of the video surveillance device generating image data based on image frames of the captured video; the processor inputting the image data into a detection model to generate zero or more detection results, wherein the detection model is composed of the convolutional neural network; the processor inputting the image data into zero or more tracking models to generate zero or more tracking results, wherein the tracking model uses a portion of the convolutional neural network; and the processor selecting targets from the detection results that have a detection threshold (T) higher than a first detection threshold (T). l The processor selects zero or more target detection results from the tracking results that have a first tracking threshold (T). corr The processor performs a detection score enhancement operation on zero or more target tracking results of the first tracking score; to generate enhanced detection results based on the number of target detection results and the tracking results; and the processor selects targets with a detection score higher than a second detection threshold (T) from the enhanced detection results. m The second detection score of the target enhancement detection results; the processor performs a matching operation on the target enhancement detection results and the target tracking results to generate a matching output, wherein the matching output includes zero or more matching results and zero or more non-matching target detection results and zero or more non-matching target tracking results, wherein each matching result has a pair of matching target enhancement detection results and target tracking results, wherein the target object within the displayed image frame is marked according to the generated matching output.
[0006] According to another aspect of the present invention, a video surveillance device is provided for detecting and tracking target objects in captured video using a convolutional neural network, and the video surveillance device includes one or more processors configured to execute machine instructions to implement the above-described method.
[0007] According to another aspect of the invention, a system is provided for detecting and tracking target objects in video captured by a video surveillance device using a convolutional neural network, and the server of the system includes one or more processors configured to execute machine instructions to implement the methods described above. Attached Figure Description
[0008] The embodiments of the present invention will now be described in more detail with reference to the accompanying drawings, wherein:
[0009] Figure 1 A block diagram of a video surveillance device according to an embodiment of the present invention is depicted;
[0010] Figure 2 A block diagram of a system according to an embodiment of the present invention is depicted;
[0011] Figure 3A A flowchart is depicted for detecting and tracking target objects in the captured video;
[0012] Figure 3B A schematic diagram of the structure of the convolutional neural network used in the detection and tracking models is shown.
[0013] Figure 3C A schematic diagram depicts the tracking initialization and tracking prediction performed by the tracking model;
[0014] Figure 3D A schematic diagram depicts the initialization of important feature selection (IFS) and the execution of IFS.
[0015] Figure 4 A schematic diagram of detection score enhancement (DSE) is depicted.
[0016] Figure 5A Depicting Figure 3A The flowchart of step S380 in the process;
[0017] Figure 5B Depicting Figure 3A The flowchart of step S390 in the text;
[0018] Figure 5C Depicting Figure 3A A further flowchart of step S390 in the process;
[0019] Figure 6 A schematic diagram illustrating the operation flow of the provided method is shown; and
[0020] Figure 7 An example of detecting and tracking a target object is described. Detailed Implementation
[0021] In the following description, methods, electronic devices, and systems for using convolutional neural networks (CNNs) to detect and track target objects in video endoscopy are listed as preferred examples. It will be apparent to those skilled in the art that modifications, including additions and / or substitutions, can be made without departing from the scope and spirit of the invention. Specific details may be omitted so as not to obscure the invention; however, this disclosure is prepared to enable those skilled in the art to practice the teachings herein without excessive experimentation.
[0022] Please refer to the following description. Figure 1 According to various embodiments of the present invention, a video surveillance device 100 for using a convolutional neural network to detect and track target objects in captured video includes a processor 110, a data communication circuit 120, a non-transient storage circuit 130, and a camera 140. In embodiments, the video surveillance device 100 may be an electronic device, such as a video endoscope, a drone with a camera, or a traffic monitoring camera.
[0023] Data communication circuit 120 is configured to establish a network connection with other electronic devices (i.e., cloud servers or backend servers). Video surveillance device 100 can receive control data CD or object data OD from other electronic devices through the established network connection. Control data CD may include data used to train tracking and detection models, data from the trained detection model, data of the determined detection / tracking results, and auxiliary data. For example, object data OD is image or video data, including multiple image frames input to video surveillance device 100 to detect and track potential target objects in the image frames.
[0024] Camera 140 is configured to capture (shoot) images / videos, generate image data (object data OD), and transmit it to processor 110.
[0025] Input / output (I / O) circuitry 150 is connected to a touchscreen or other suitable image / video display device via wired or wireless means. In one embodiment, processor 110 analyzes the object data OD to obtain result data and instructs I / O circuitry 150 to transmit the display data signal to display image frames and markers corresponding to the target object based on the result data.
[0026] On the other hand, a system is provided that utilizes convolutional neural networks to detect and track target objects in captured videos. For example... Figure 2 As shown, system 1 includes a video surveillance device 100 and a server 200. The server 200 includes a processor 210, a data communication circuit 220, and a non-transient storage circuit 230. The data communication circuit 220 is configured to establish a network connection NC with the video surveillance device 100.
[0027] The non-transient storage circuits 130 / 230 are configured to store programs 131 / 231 (or machine instructions 131 / 231) and managed databases 132 / 232. Databases 132 / 232 can be used to store trained detection models (also known as detectors), tracking models (also known as trackers), object data (OD), control data (CD), and / or analysis results (such as generated detection and tracking results, also known as result data (RD)).
[0028] Processors 110 / 210 are used to execute machine instructions 131 / 231 to implement the methods provided in this disclosure. The aforementioned detection model and tracking model are executed by processors 110 / 210.
[0029] In one embodiment, server 200 analyzes the received object data OD and sends the result data RD to another electronic device 300 so that a tag corresponding to the target object can be displayed based on the result data RD. The electronic device 300 may be, for example, a computer, a surveillance camera, etc.
[0030] A target object is an image object that an electronic server searches for, locates, and marks within an image frame of image data. For example, in the field of video endoscopy, the target object is a lesion in the video frame; in traffic monitoring, the target object can be a vehicle, pedestrian, or other types of moving objects.
[0031] The video surveillance device 100 is an exemplary embodiment used to illustrate the provided method.
[0032] refer to Figure 3A In step S300, the camera 140 generates image data (e.g., target data) based on the image frames of the captured video. The image data is then sent to the processor 110.
[0033] In step S310, the processor 110 inputs image data into a detection model to generate zero or more detection results. The step of inputting image data into the detection model to generate detection results includes: inputting the image data into a convolutional neural network to obtain one or more features of the image frame; determining zero or more detection marker positions, detection scores, and zero or more target object types based on the features; generating detection results based on the detection marker positions and target object types, wherein each detection result includes a corresponding detection marker position and corresponding label information, where the label information includes the target object type of the corresponding detection result. The target object type can be determined from preset object types. The processor 110 can instruct the I / O circuit 150 to display the target object type next to the detection marker based on the label information.
[0034] Specifically, such as Figure 3B As shown, the neural network structure of the detection model contains M convolutional blocks. In a preferred embodiment, M is 5. Each convolutional block includes one or more convolutional layers and one or more activation functions. In a preferred embodiment, residual blocks are used.
[0035] The M convolutional blocks and the detection model are trained on a labeled dataset of target objects, thus ensuring that the trained convolutional blocks have strong representational power to describe the target objects.
[0036] In addition, refer to Figure 3B and Figure 3C The tracking model includes a Dedicated Feature Extractor (DFE) 310 and an Important Features Selection (IFS) 320. The DFE is constructed from the first three convolutional blocks out of M convolutional blocks. The DFE extracts multi-resolution high-dimensional features (e.g., dedicated features) from the input image data (e.g., image frames) and outputs them to the IFS 320. Leveraging the strong feature representation capability of the DFE, the IFS can directly select N% of the highly activated features as the most useful features (e.g., typical dedicated features) using spatial averaging. N can be, for example, 10 or other values, and this invention is not limited to this. The DFE 310 works in conjunction with the IFS 320 to improve tracking speed and reduce processing latency, thereby ensuring real-time processing.
[0037] More specifically, such as Figure 3C The upper part (Tracking Initialization) shows that when the detection model generates new detection results, the processor 110 will build and initialize a new tracking model. Qualified detection results (when their scores are higher than the threshold T) will be used. h The target object image corresponding to the given time is input into the DFE 310 (part of the convolutional neural network) to obtain the corresponding dedicated features. These dedicated features are then input into the IFS initialization 321 to obtain typical indices, which are recorded. The processor 110 selects typical dedicated features from the dedicated features DF based on the recorded typical indices. These typical dedicated features are then used to build and initialize a tracking model based on a discriminative correlation filter (DCF). Since the tracking model learns correlation filters from the target's appearance features to distinguish between the target and the background appearance, it is called a tracking model based on a discriminative correlation filter (DCF).
[0038] More specifically, during the tracking initialization process, a DCF-based tracking model is established and initialized. This step includes: inputting typical specific features of the target object image; transforming the typical specific features to the frequency domain using Fast Fourier Transform (FFT); generating the transformed typical specific features; and training a correlation filter using the transformed typical specific features in the frequency domain to distinguish the appearance of the target and the background.
[0039] After establishing and initializing the tracking model (tracking initialization), processor 110 will continue to use the tracking model to perform tracking predictions for subsequent video frames until the tracking model is removed. For example... Figure 3C As shown in the lower half (Tracking Prediction), the search region image is input to DFE 310. Processor 110 determines a portion of the image frame from the previous frame that surrounds the target position of the tracking result in the previous frame as the search region image. DFE 310 can then generate specific features from the search region image. "IFS Execution" 322 reads the recorded typical index to obtain typical specific features from the specific features. Finally, the initialized DCF-based tracking model uses the typical specific features to predict the target position of the target object as the tracking result TR.
[0040] The following will be passed Figure 3D Let me explain the IFS process in detail.
[0041] Please refer to Figure 3D During IFS initialization, after the target object image corresponding to the detection result is input into a part of the convolutional neural network to obtain the dedicated features DF, the processor 110 performs global average pooling on the dedicated features DF to obtain average features AF. Then, the processor 100 sorts the average features AF in descending order to obtain sorted features SF and an index array (IDX) of sorted features SF. Next, the processor 110 selects the top N% of the index array of sorted features as typical indices TIDX, which are recorded. In other words, IFS helps select the top N% of important features to reduce the dimensionality of dedicated features, reduce the computational cost of tracking prediction, thereby reducing the time and cost of object tracking and improving tracking efficiency.
[0042] For example, let's assume the average feature is [3,6,4,1], with indices [0,1,2,3]. In this example, the sorted feature is [6,4,3,1], and the index array of the sorted feature is [1,2,0,3]. If we choose the first two (e.g., N=50), the typical index would be [1,2].
[0043] Furthermore, during IFS execution, processor 110 inputs the search region image into a portion of a convolutional neural network to obtain specialized features (DFs). The location of the search region can be determined, for example, based on the tracking results of the previous frame. Processor 110 accesses a recorded typical index and selects a typical specialized feature (TDF) from the specialized features (DFs) based on the recorded typical index (IDX). In other words, the typical index is recorded during IFS initialization and used during IFS execution.
[0044] Refer again Figure 3A In step S330, the processor 110 selects from the detection results those with values higher than the first detection threshold (T). l The processor 110 determines whether the first detection score of the detection result is higher than a first detection threshold (T). l If the first detection score of the detection result is higher than the first detection threshold, the processor 110 performs a detection score enhancement (DSE) operation based on the detection result and the target tracking result to obtain an enhanced detection result (step S350).
[0045] Artifacts are very common in endoscopic videos, producing many low-quality frames. Examples include overexposure / underexposure, motion blur, video defocus, fluids, bubbles, specular reflections, and floating objects. Since the detection model is trained on qualified training images, directly applying it to these low-quality frames will produce low-confidence detection results (the model may find the target object, but its detection score is low). The purpose of Disturbance Sequence (DSE) is to utilize the temporal information of the image to compensate for image quality defects, thereby helping to enhance some low-confidence detection results and ultimately improving detection accuracy.
[0046] Specifically, given a detection result d, find a tracking model t whose tracking result has the maximum overlap with d (measured by intersection over union, IoU), where q is the last associated detection result of tracking model t. The enhanced detection result is then scored. It can be represented by the following formula (1).
[0047]
[0048] Among them, Y(d) It is the score of the test result d, Y (q) M is the score of the last association detection result of the tracking model t. (t) (Match count) refers to the number of consecutive frames whose detection results are associated with the tracking model t. (t) (Mismatch count) refers to the number of consecutive frames in which the detection results are not associated with the tracking model t. λ is the confidence parameter for long-term detection, and β is the uncertainty parameter for consecutive mismatches (default λ = 2, β = 1.5).
[0049] like Figure 4 As shown, assuming image frame IF1 is input to the detection model, a detection result DR1 with a detection score of 32% is generated. This detection result is associated with a certain tracking model. Then, image frame IF2 following image frame IF1 is input to the detection model and the tracking model, resulting in a detection result DR2 with a detection score of 10% and a tracking result TR. Processor 110 performs a detection score enhancement operation according to formula (1) to obtain an enhanced detection result DR3 for image frame IF2', which has an enhanced detection score of 28% higher than the original detection score of 10%. It should be noted that "003,cancer,28%" is the label displayed based on the label information of the enhanced detection result DR3. "003" is the ID of the target object, and "cancer" is the type of the target object.
[0050] Refer again Figure 3A In step S360, the processor 110 determines whether the second detection score of the enhanced detection result is higher than the second detection threshold (T). m The processor 110 selects an enhanced detection result with a score higher than the second detection threshold to perform step S370.
[0051] Furthermore, in step S320, the processor 110 inputs image data into zero or more tracking models to generate zero or more tracking results. Each tracking result has its own tracking model. The step of inputting image data into zero or more tracking models to generate zero or more tracking results includes: inputting a search region of the image frame into a portion of a convolutional neural network to obtain second specific features; accessing recorded typical indices; selecting typical specific features from the second specific features based on the recorded typical indices; inputting the typical specific features into each tracking model to predict the target position and output a response score (i.e., a tracking score), wherein the tracking model is a DCF-based tracking model; determining the tracking marker position based on the predicted target position; and generating the tracking result based on the tracking marker position, the tracking result including the tracking marker position and the tracking score. Typical specific features are input into a tracking model based on a discriminative correlation filter (DCF), and the tracking model outputs the predicted target position and its response score (i.e., the tracking score).
[0052] More specifically, the tracking prediction includes: inputting typical specific features of the search region image; transforming the typical specific features to the frequency domain using Fast Fourier Transform (FFT); generating the transformed typical specific features; calculating the Fourier response map by performing element-wise multiplication of the trained correlation filter with the transformed typical specific features in the frequency domain; summing and adding the Fourier response maps of multiple typical specific features to generate a summed Fourier response map; transforming the summed Fourier response map to the spatial domain using inverse Fourier transform to generate a spatial response map; identifying the location with the largest response value from the spatial response map; outputting the identified location as the new target location, and outputting the maximum response value as the tracking score.
[0053] Next, in step S340, the processor 110 selects from the tracking results items that have a value higher than the first tracking threshold (T). corr The processor 110 determines whether the first tracking score of the tracking result is higher than the first tracking threshold (T). corr The processor 110 selects a tracking result with a score higher than the first tracking threshold to perform step S370.
[0054] In step S370, the processor 110 performs a matching operation on the target enhancement detection results and the target tracking results to generate a matching output. The matching output includes zero or more matching results, zero or more non-matching target detection results, and zero or more non-matching target tracking results, where each matching result has a pair of matching target enhancement detection results and target tracking results. The target objects within the displayed image frame are marked according to the generated matching output (as in steps S380 and S390). For example, assuming there are X target enhancement detection results and Y target tracking results, the matching operation will produce Z matching results, XZ non-matching target detection results, and YZ non-matching target tracking results.
[0055] The matching operation employs the Hungarian Algorithm. Specifically, the Hungarian Algorithm is used to match the target augmentation detection results with the target tracking results, where the Intersection over Union (IoU) between each detection box (detection result) and the tracked box (tracking result) is calculated as the allocation cost. An IoU threshold of 0.2 is used to filter out matching pairs with low overlap.
[0056] In step S380, processor 110 processes the matching result. In step S390, processor 110 processes the non-matching target detection result and the non-matching target tracking result.
[0057] refer to Figure 5A In step S381, for each matching result, a pair of matching target enhancement detection results and target tracking results, the processor 110 identifies the target tracking model that generates the target tracking result and associates the target enhancement detection result with the target tracking model. Next, in step S382, the processor 110 instructs the I / O circuit 150 to display detection markers (such as...) in the displayed image frame based on the target enhancement detection results. Figure 4 The detection markers displayed in DR1 indicate the target objects within the image frame, and the target enhancement detection result includes the detection marker positions and marker information corresponding to the target objects.
[0058] In addition, in step S383, the processor 110 performs a tracking reliability estimation to obtain a reliability score for the corresponding target tracking result.
[0059] Given a tracking model t that produces tracking results, its tracking reliability is estimated by the last associated detection result q of the tracking model t, as shown in the following formula (2).
[0060]
[0061] Among them, Z (t)It is the score of the tracking result, and also the current tracking score of the tracking model t, Y. (q) It is the detection score of its last associated detection result q, U (t) (Mismatch count) refers to the number of consecutive frames in which the detection results are not associated with the tracking model t. α is the uncertainty parameter for consecutive mismatches (default is α = 0.1).
[0062] The detection score represents the target object's aimingness, while the tracking score is the correlation response between the tracked target object and the detection results of several previous frames. The product of these scores describes the target object's aimingness and reflects the reliability of the current tracking result.
[0063] Tracking reliability estimation (TRE) is helpful for tracking models and their tracking results: if a matching detection result is found, the tracking model is updated when the corresponding TRE score (also known as the reliability score) is higher than a given threshold; otherwise, if no matching detection result is found, the tracking result will be generated when the TRE is greater than the given threshold.
[0064] Selectively updating a tracking model with a high reliability score can remove any unreliable training samples to avoid tracking drift, thereby improving tracking robustness. Selectively using tracking results with high reliability scores for targets missed by the detection model can produce more stable tracking, thus improving monitoring visualization.
[0065] The steps for updating the tracking model include: (a) inputting the target object image corresponding to the tracking result into a part of the convolutional neural network to obtain the dedicated feature DF; (b) inputting the dedicated feature DF into the IFS for execution to obtain the typical dedicated feature TDF; (c) adding the typical dedicated feature TDF as a new training sample; (d) when the number of new samples exceeds K (K=10), using all training samples to train the tracking model, and resetting the counter of the new samples after training.
[0066] Next, in step S384, if the reliability score is higher than the reliability threshold (T) rel The processor 110 updates the target tracking model based on the target tracking result. Otherwise, if the reliability score is not higher than the reliability threshold (T)... rel The processor 110 will not update the tracking model that generated the target tracking result.
[0067] Furthermore, some control parameters are updated when the enhanced detection results are correlated with the tracking model t. For example, (1) ifU (t) >0:U (t) =0,M (t) =0; (2)M (t) + = 1.
[0068] Reference Figure 5B (Processing mismatched tracking results), in step S391, the processor 110 identifies the target tracking model that generated the mismatched target tracking result for each mismatched target tracking result, and updates the mismatch count (U) of the target tracking model. (t) Some control parameters will be updated, for example, U. (t) + = 1, where t is the tracking model.
[0069] Next, in step S392, the processor 110 determines the mismatch count (U) of the target tracking model. (t) Is it higher than the mismatch count threshold (U)? TH ).
[0070] If the count does not match (U) (t) If the mismatch count exceeds the mismatch count threshold, in step S393, the processor 110 removes the target tracking model. (t) If the result is not higher than the mismatch count threshold, then in step S394, the processor 110 performs a tracking reliability estimation to obtain a reliability score corresponding to the target tracking result.
[0071] Next, in step S395, if the reliability score is higher than the reliability threshold (T) rel The processor 110 instructs the I / O circuit 150 to display tracking markers in the displayed image frame based on the generated target tracking results. The displayed tracking markers indicate the target object within the image frame, and the target tracking results include the tracking marker positions. If the reliability score is not higher than the reliability threshold (T...), rel The processor 110 will ignore this target tracking result.
[0072] On the other hand, refer to Figure 5C (Processing mismatch detection results), in step S396, the processor 110 enhances the detection result for each mismatched target and determines whether the second detection score of the target enhancement detection result is higher than the third detection threshold (T). h If the second detection score is not higher than the third detection threshold (T) h In step S397, the processor 110 ignores the enhanced detection result of the mismatched target; otherwise, if the second detection score is higher than the third detection threshold (T... hFollowing step S398, the processor instructs the I / O circuit 150 to display another detection marker within the displayed image frame based on the target enhancement detection result. The displayed other detection marker indicates another target object within the image frame. The enhancement detection result includes the location of another detection marker corresponding to the other target object and other marker information.
[0073] Next, in step S399, if the second detection score of the target enhancement detection result is higher than the third detection threshold (T) h The processor 110 uses the target enhancement detection results to establish and initialize a new tracking model. The steps of establishing and initializing the new tracking model using the target detection results include: inputting the target object image corresponding to the target detection results into a part of a convolutional neural network to obtain first specific features; performing global average pooling on the first specific features to obtain average features AF; sorting the average features in descending order to obtain sorted features SF, and obtaining an index array IDX of the sorted features; selecting the top N% of the index array as typical indices TIDX, where typical indices are recorded; selecting typical specific features TDF from the first specific features based on the recorded typical indices; and establishing and initializing a tracking model based on a discriminative correlation filter (DCF) using the typical specific features.
[0074] Next, in step S400, the processor 110 associates the target enhancement detection result with the new tracking model. This is the first association established between the new tracking model and the detection result. Furthermore, some control parameters are updated during this initial association establishment. For example, U (t) =0,M (t) =1, where t is the tracking model.
[0075] like Figure 6 As shown, image frames are input into the detection model (610) to obtain detection results (611), and into the tracking model (620) to obtain tracking results (621). A score is assigned to each detection result and compared with a threshold T. l The results are compared (612). A score for each tracking result is determined and compared with a threshold T. corr Compare (622). When the detection score is higher than the threshold T l At that time, detection scoring enhancement is performed (613). The score of each enhanced detection result is determined and compared with the threshold T. m Comparison (614). Matching is performed by executing a matching operation using the Hungarian algorithm to match scores higher than the threshold T. m Enhanced detection results and scores higher than the threshold T corrThe tracking results (630) are used to calculate the overlap (IoU) between each detection box (detection marker) and the tracking box (tracking marker) as the assignment cost. l : A low scoring threshold used to retrieve all possible detection results. T m : A scoring threshold used to retrieve candidate detection results that can be correlated with the tracking results. T corr : A threshold used to filter tracking results with very low relevance response.
[0076] For each matching detection and tracking result, the control parameters of the tracking model that generated the tracking result are updated, and the enhanced detection result is associated with the tracking model (641). Then, tracking reliability estimation (TRE) is performed to obtain a TRE score (642), which is then compared with a threshold T. rel Comparison (643). When the TRE score is higher than the threshold T rel At this time, the tracking model that generates the tracking results will be updated (644). In addition, all matching detection results will be generated (645). T rel : Threshold used to select a tracking model with high tracking reliability.
[0077] For mismatch tracking results, the corresponding control parameters (such as the mismatch count U) (t) The tracking reliability estimate (TRE) will be updated (651). Then, a tracking reliability score (652) is obtained by performing a tracking reliability estimate (TRE) and comparing the score with a threshold T. rel Comparison. When the TRE score is higher than the threshold T... rel At that time, a tracking result (653) is generated. It should be noted that when the mismatch count exceeds the threshold U... TH When this happens, the tracking model will be removed / disabled (654).
[0078] The score is higher than the threshold T h The mismatch detection results will be generated (661). For each mismatch detection result generated, a new tracking model will be added, the control parameters of the newly established tracking model will be updated, and the mismatch detection result will be associated with the newly established tracking model (662). h : The high-score threshold used to retrieve high-confidence detection results.
[0079] Please refer to Figure 7 For example, at time T1, image frame IF1 is input into the detection model to obtain the detection result. Assume the detection result score is higher than the threshold T. hA tracking model is built using the detection results. The search region SA1 is determined based on the detection results. The tracking model will use the search region SA1 to track the target object in the next image frame (e.g., at time T2). The detection marker DM1_1 will be displayed. Simultaneously, based on the label information from the detection results, the corresponding label "Target Object #1_1" will be displayed near the detection marker.
[0080] At time T2, image frame IF2 is input into the detection model to obtain another detection result. The search region SA1 of image frame IF2 is input into the established tracking model to obtain a tracking result. Assuming the other detection result matches the tracking result, the detection marker DM1_2 is displayed. Simultaneously, based on the label information of the detection result, the corresponding label "Target Object #1_1" is displayed near the detection marker. Another search region SA2 is determined based on the current tracking result, and search region SA2 will be used by the tracking model to track the target object in the next image frame (e.g., at time T3).
[0081] At time T3, image frame IF3 is input into the detection model to obtain another detection result. Assuming the other detection result matches the tracking result, the detection marker DM1_3 is displayed. Simultaneously, based on the label information of the detection result, the corresponding label "Target Object #1_1" is displayed near the detection marker.
[0082] At time T4, image frame IF4 is input into the detection model to obtain another detection result. Assume that this other detection result does not match the tracking result, and the TRE score of the mismatched tracking result is higher than that of T. rel If the score is higher than the threshold T, then the tracking marker TM1 will be displayed. h A new tracking model is then built using the mismatch detection results. Therefore, there are now two tracking models. Detection marker DM1_4 will be displayed. Furthermore, based on the label information of the detection results, the corresponding label "Target Object #1_2" will be displayed near the detection marker.
[0083] At time T5, image frame IF5 is input into the detection model to obtain another detection result, and then input into two tracking models to obtain two tracking results. Assuming the other detection result matches one of the tracking results, a detection marker DM1_5 is displayed. Based on the label information of the detection result, the corresponding label "Target Object #1_2" is displayed near the detection marker. For the other non-matching tracking result, it is assumed that the non-match count of the tracking model that generated this tracking result is greater than a threshold U. TH If this happens, the tracking model will be removed. Therefore, only one tracking model remains.
[0084] At time T6, image frame IF6 is input to the detection model, yielding multiple additional detection results. If one of these detection results matches the tracking result, the corresponding detection marker DM1_6 is displayed, along with the label "Target Object #1_2". Furthermore, if the score of another non-matching detection result is higher than a threshold T... h A new tracking model is built using the mismatch detection results. The detection marker DM2_1 is displayed. Simultaneously, the corresponding label "Target Object #2_1" is also displayed.
[0085] The above exemplary implementations and operations are merely illustrative of the present invention. Those skilled in the art will understand that other structural and functional configurations and applications are possible and readily adopted without unnecessary experimentation and deviation from the spirit of the invention.
[0086] The functional units of the devices and methods according to the embodiments disclosed herein may be implemented using computing devices, computer processors, or electronic circuit systems, including but not limited to application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and other programmable logic devices configured or programmed according to the teachings of this disclosure. Machine instructions or firmware / software code running in the computing device, computer processor, or programmable logic device can be readily created by those skilled in the art based on the teachings of this disclosure.
[0087] All or part of the methods according to the embodiments can be executed in one or more computing devices including server computers, personal computers, laptop computers, mobile computing devices (e.g., smartphones) and tablet computers.
[0088] The embodiments include non-transitory memory circuitry and / or computer storage media having machine instructions or firmware / software code stored therein, which can be used to program a processor to perform any of the processes of the present invention. The non-transitory memory circuitry and / or storage media include (but are not limited to) floppy disks, optical disks, Blu-ray discs, DVDs, CD-ROMs and magneto-optical disks, ROMs, RAMs, flash memory devices, or any type of medium or device suitable for storing instructions, code, and / or data.
[0089] Each of the functional units according to the various embodiments can also be implemented in a distributed computing environment and / or cloud computing environment, wherein one or more processing devices interconnected via communication networks such as intranets, wide area networks (WANs), local area networks (LANs), the Internet, and other forms of data transmission media execute all or part of machine instructions in a distributed manner. The communication networks established in the various embodiments support various communication protocols, such as (but not limited to) Wi-Fi, Global System for Mobile Communications (GSM) systems, Personal Handheld Phone Systems (PHS), Code Division Multiple Access (CDMA) systems, Global Microwave Access Interoperability (WiMAX) systems, third-generation wireless communication technology (3G), fourth-generation wireless communication technology (4G), fifth-generation wireless communication technology (5G), Long Term Evolution (LTE), Bluetooth, and Ultra Wideband (UWB).
[0090] The foregoing description of the invention has been provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations will be apparent to those skilled in the art.
[0091] The embodiments were selected and described in order to best illustrate the principles of the invention and its practical application, thereby enabling others skilled in the art to understand the various embodiments of the invention and the various modifications suitable for particular uses upon careful consideration.
Claims
1. A computer-based method for detecting and tracking target objects in captured video using a convolutional neural network (CNN) via video surveillance equipment, characterized in that... The method includes: The processor of the video surveillance equipment generates image data based on the image frames of the captured video. The processor inputs the image data into the detection model to generate zero or more detection results, wherein the detection model is composed of the convolutional neural network (S310). The processor inputs the image data into zero or more tracking models to generate zero or more tracking results, wherein the tracking models use a portion of the convolutional neural network (S320). The processor selects from the detection results those with values higher than a first detection threshold (T). l The first detection score of zero or more target detection results (S330); The processor selects from the tracking results items that have a value higher than a first tracking threshold (T). corr The first tracking score of zero or more target tracking results (S340); The processor performs a detection scoring enhancement operation to generate an enhanced detection result based on the number of target detection results and the target tracking results (S350); The processor selects from the enhanced detection results those with values higher than the second detection threshold (T). m The second detection score includes zero or more target enhancement detection results (S360); The processor performs a matching operation on the target enhancement detection result and the target tracking result to generate a matching output (S370), wherein the matching output includes zero or more matching results and zero or more non-matching target detection results and zero or more non-matching target tracking results, wherein each matching result has a pair of matched target enhancement detection results and target tracking results. The target object within the displayed image frame is marked according to the generated matching output.
2. The computer implementation method according to claim 1, further comprising processing the matching result (S380), characterized in that, The steps for processing the matching results include: For each matching result, the enhanced detection result of the matched target and the target tracking result, Identify the target tracking model that generated the target tracking results; The target enhancement detection results are correlated with the target tracking model; Perform a tracking reliability estimate to obtain a reliability score corresponding to the target tracking result; When the reliability score is greater than the reliability threshold (T) rel When the target tracking result is obtained, the target tracking model is updated. as well as Based on the target enhancement detection result, a detection marker is displayed in the displayed image frame. The displayed detection marker indicates the target object within the image frame. The target enhancement detection result includes the detection marker position and label information corresponding to the target object.
3. The computer implementation method according to claim 1, characterized in that, The process further includes processing the mismatched target detection results, wherein the steps for processing the mismatched target detection results include: For each mismatched target, enhance the detection result: Determine whether the second detection score of the target enhancement detection result is higher than the third detection threshold (T). h ); If the second detection score of the target enhancement detection result is not higher than the third detection threshold (T) h If the target mismatch is detected, the enhanced detection result is ignored. If the second detection score of the target enhancement detection result is higher than the third detection threshold (T) h If the target enhancement detection result is obtained, another detection marker is displayed in the displayed image frame, wherein the displayed other detection marker indicates another target object within the image frame, and the target enhancement detection result includes the location of another detection marker corresponding to the other target object and another label information; A new tracking model is established and initialized using the target augmentation detection results; and The target enhancement detection results are correlated with the new tracking model.
4. The computer implementation method according to claim 3, characterized in that, The process further includes processing mismatched target tracking results, wherein the steps for processing the mismatched target tracking results include: For each non-matching target tracking result: Identify the target tracking model that produces the mismatched target tracking results; Determine the mismatch count of the target tracking model; The mismatch count (U) of the target tracking model is determined. (t) Is it higher than the mismatch count threshold (U)? TH ); If the mismatch count of the target tracking model is higher than the mismatch count threshold, then the target tracking model is removed. as well as If the mismatch count of the target tracking model is not higher than the mismatch count threshold, Perform a tracking reliability estimate to obtain a reliability score corresponding to the target tracking result; as well as When the reliability score is greater than the reliability threshold (T) rel When the target tracking result is generated, a tracking marker is displayed in the displayed image frame, wherein the displayed tracking marker indicates the target object within the image frame, and wherein the target tracking result includes the tracking marker position.
5. The computer implementation method according to claim 3, characterized in that, The steps for establishing and initializing the new tracking model using the target detection results include: The target object image corresponding to the target detection result is input into the part of the convolutional neural network to obtain the first specific feature; The first dedicated feature is subjected to global average pooling to obtain the average feature (AF); The average features are sorted in descending order to obtain sorted features (SF), and an index array (IDX) of the sorted features is obtained. The top N% of the index array are selected as typical indices (TIDX), and these typical indices are recorded. Based on the recorded typical index, select typical specific features (TDFs) from the first specific features; and Using the aforementioned typical specific features, a tracking model based on the Discriminative Correlation Filter (DCF) is established and initialized.
6. The computer implementation method according to claim 1, characterized in that, The step of inputting the image data into the tracking model to generate the tracking result includes: For each tracking model: The search region of the image frame is input into the part of the convolutional neural network to obtain the second specific feature; Access the recorded typical indexes; Based on the recorded typical index, select typical special features from the second special features; The typical specific features are input into the tracking model to predict the target position within the image frame, wherein the tracking model is a DCF-based tracking model; as well as The tracking marker position is determined based on the predicted target position; as well as The tracking result is generated based on the tracking marker position, wherein the tracking result includes the corresponding tracking marker position.
7. The computer implementation method according to claim 1, characterized in that, The step of inputting the image data into the detection model to generate the detection result includes: The image data is input into the convolutional neural network to obtain the features of the image frame; Based on the aforementioned features, determine zero or more detection marker locations, detection scores, and the types of zero or more target objects; The detection result is generated based on the detection mark position and the type of the target object, wherein each detection result includes the corresponding detection mark position and the corresponding label information, wherein the label information includes the type of the target object corresponding to the detection result.
8. A video surveillance device that uses a convolutional neural network (CNN) to detect and track target objects in captured video, characterized in that, The video surveillance device includes: Camera, configured to capture video; and A processor configured to execute machine instructions to implement a method for detecting and tracking the target object, the method comprising: The processor of the video surveillance equipment generates image data based on the image frames of the captured video. The processor inputs the image data into the detection model to generate zero or more detection results, wherein the detection model is composed of the convolutional neural network (S310). The processor inputs the image data into zero or more tracking models to generate zero or more tracking results, wherein the tracking models use a portion of the convolutional neural network (S320). The processor selects from the detection results those with values higher than a first detection threshold (T). l The first detection score of zero or more target detection results (S330); The processor selects from the tracking results items that have a value higher than a first tracking threshold (T). corr The first tracking score of zero or more target tracking results (S340); The processor performs a detection scoring enhancement operation to generate an enhanced detection result based on the number of target detection results and the target tracking results (S350); The processor selects from the enhanced detection results those with values higher than the second detection threshold (T). m The second detection score includes zero or more target enhancement detection results (S360); The processor performs a matching operation on the target enhancement detection result and the target tracking result to generate a matching output (S370), wherein the matching output includes zero or more matching results and zero or more non-matching target detection results and zero or more non-matching target tracking results, wherein each matching result has a pair of matched target enhancement detection results and target tracking results. The target object within the displayed image frame is marked according to the generated matching output.
9. The video surveillance device according to claim 8, further comprising processing the matching result (S380), characterized in that, The steps for processing the matching results include: For each matching result, the enhanced detection result of the matched target and the target tracking result, Identify the target tracking model that generated the target tracking results; The target enhancement detection results are correlated with the target tracking model; Perform a tracking reliability estimate to obtain a reliability score corresponding to the target tracking result; When the reliability score is greater than the reliability threshold (T) rel When the target tracking result is obtained, the target tracking model is updated. as well as Based on the target enhancement detection result, a detection marker is displayed in the displayed image frame. The displayed detection marker indicates the target object within the image frame. The target enhancement detection result includes the detection marker position and label information corresponding to the target object.
10. The video surveillance device according to claim 8, characterized in that, The process further includes processing the mismatched target detection results, wherein the steps for processing the mismatched target detection results include: For each mismatched target, enhance the detection result: Determine whether the second detection score of the target enhancement detection result is higher than the third detection threshold (T). h ); If the second detection score of the target enhancement detection result is not higher than the third detection threshold (T) h If the target mismatch is detected, the enhanced detection result is ignored. If the second detection score of the target enhancement detection result is higher than the third detection threshold (T) h If the target enhancement detection result is obtained, another detection marker is displayed in the displayed image frame, wherein the displayed other detection marker indicates another target object within the image frame, and the target enhancement detection result includes the location of another detection marker corresponding to the other target object and another label information; A new tracking model is established and initialized using the target augmentation detection results; and The target enhancement detection results are correlated with the new tracking model.
11. The video surveillance device according to claim 10, characterized in that, The process further includes processing mismatched target tracking results, wherein the steps for processing the mismatched target tracking results include: For each non-matching target tracking result: Identify the target tracking model that produces the mismatched target tracking results; Determine the mismatch count of the target tracking model; The mismatch count (U) of the target tracking model is determined. (t) Is it higher than the mismatch count threshold (U)? TH ); If the mismatch count of the target tracking model is higher than the mismatch count threshold, then the target tracking model is removed. as well as If the mismatch count of the target tracking model is not higher than the mismatch count threshold, Perform a tracking reliability estimate to obtain a reliability score corresponding to the target tracking result; as well as When the reliability score is greater than the reliability threshold (T) rel When the target tracking result is generated, a tracking marker is displayed in the displayed image frame, wherein the displayed tracking marker indicates the target object within the image frame, and wherein the target tracking result includes the tracking marker position.
12. The video surveillance device according to claim 10, characterized in that, The steps for establishing and initializing the new tracking model using the target detection results include: The target object image corresponding to the target detection result is input into the part of the convolutional neural network to obtain the first specific feature; The first dedicated feature is subjected to global average pooling to obtain the average feature (AF); The average features are sorted in descending order to obtain sorted features (SF), and an index array (IDX) of the sorted features is obtained. The top N% of the index array are selected as typical indices (TIDX), and these typical indices are recorded. Based on the recorded typical index, select typical specific features (TDFs) from the first specific features; and Using the aforementioned typical specific features, a tracking model based on the Discriminative Correlation Filter (DCF) is established and initialized.
13. The video surveillance device according to claim 8, characterized in that, The step of inputting the image data into the tracking model to generate the tracking result includes: For each tracking model: The search region of the image frame is input into the part of the convolutional neural network to obtain the second specific feature; Access the recorded typical indexes; Based on the recorded typical index, select typical special features from the second special features; The typical specific features are input into the tracking model to predict the target position within the image frame, wherein the tracking model is a DCF-based tracking model; The tracking marker position is determined based on the predicted target position; as well as The tracking result is generated based on the tracking marker position, wherein the tracking result includes the corresponding tracking marker position.
14. The video surveillance device according to claim 8, characterized in that, The step of inputting the image data into the detection model to generate the detection result includes: The image data is input into the convolutional neural network to obtain the features of the image frame; Based on the aforementioned features, determine zero or more detection marker locations, detection scores, and the types of zero or more target objects; The detection result is generated based on the detection mark position and the type of the target object, wherein each detection result includes the corresponding detection mark position and the corresponding label information, wherein the label information includes the type of the target object corresponding to the detection result.
15. A system for detecting and tracking target objects in a captured video using a convolutional neural network (CNN), comprising: A video surveillance device; and A server, the server comprising: processor, The video surveillance device sends object data, including the captured video, through the network established between the video surveillance device and the server. The processor is configured to execute machine instructions to implement a method for detecting and tracking the target object, the method comprising: The processor of the video surveillance equipment generates image data based on the image frames of the captured video. The processor inputs the image data into the detection model to generate zero or more detection results, wherein the detection model is composed of the convolutional neural network (S310). The processor inputs the image data into zero or more tracking models to generate zero or more tracking results, wherein the tracking models use a portion of the convolutional neural network (S320). The processor selects from the detection results those with values higher than a first detection threshold (T). l The first detection score of zero or more target detection results (S330); The processor selects from the tracking results items that have a value higher than a first tracking threshold (T). corr The first tracking score of zero or more target tracking results (S340); The processor performs a detection scoring enhancement operation to generate an enhanced detection result based on the number of target detection results and the target tracking results (S350); The processor selects from the enhanced detection results those with values higher than the second detection threshold (T). m The second detection score includes zero or more target enhancement detection results (S360); The processor performs a matching operation on the target enhancement detection result and the target tracking result to generate a matching output (S370), wherein the matching output includes zero or more matching results and zero or more non-matching target detection results and zero or more non-matching target tracking results, wherein each matching result has a pair of matched target enhancement detection results and target tracking results. The target object within the displayed image frame is marked according to the generated matching output.
Citation Information
Patent Citations
End-to-End Tracking of Objects
US20190147610A1
Dynamic multi-camera tracking of moving objects in motion streams
US20200226769A1