Method and system for re-identification
By comparing object feature vectors in a video sequence and reprocessing image data, the object re-identification challenge caused by the difference in image processing settings is solved, and the accuracy and stability of object tracking are improved.
Patent Information
- Application Number
- CN202411841134.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-20
AI Technical Summary
In some settings and scenarios, there are challenges in re-identification of objects in video sequences, especially when there are significant differences between image processing settings, resulting in feature vector differences, which in turn affects the accurate tracking of objects.
By comparing the object feature vectors in the image frame, if it is found that the difference in the feature vector exceeds the threshold and the difference between the image processing settings also exceeds the threshold, the original image data is reprocessed, the updated feature vector is extracted, and the comparison is performed again to determine the similarity of the object.
This method improves the accurate re-identification of objects in different image frames, reduces false recognition due to differences in image processing settings, and enhances the stability of object tracking in video sequences.
Smart Images

Figure CN120182880A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a computer-implemented method for determining the appearance similarity of objects in image frames of at least one video sequence. The present disclosure further relates to an image processing system configured to perform the method. Background Art
[0002] Re-identification in a video sequence refers to the process of identifying and tracking objects across different image frames of the video sequence. This may be useful in many different applications for a variety of reasons. For example, in a surveillance system, it may be crucial to identify and track people or vehicles for security purposes. In other applications, it may be useful to monitor and analyze, for example, the behavior and trajectory of an object over time. Additionally, many automated systems, such as those used in autonomous vehicles or robots, rely on re-identification and object tracking to make decisions.
[0003] When tracking an object in a video depicting a scene, the object detection in the current frame is associated with the object trajectory in the previous frame. Traditionally, the association is based on comparing the state of the object detected in the current frame, such as the position, with the predicted state. More recently, methods that consider the appearance of the object being tracked have been developed.
[0004] For this purpose, feature vectors that can capture the essential features of the appearance of an object can be used for re-identification. The feature vectors of different frames can be compared or matched to determine the similarity between the objects detected in different frames. A matching algorithm can be used to evaluate the similarity between the feature vectors to determine whether they belong to the same object. One way to perform re-identification is to use machine learning to extract the feature vectors. A machine learning model can be trained on a dataset of object images to generate similar feature vectors for images depicting the same physical object and deviant feature vectors for images depicting different physical objects. Thus, a deep learning model can be used as part of the process of tracking objects in a video.
[0005] Despite the recent progress of this technology, re-identification remains challenging for certain setups and scenarios. As an example, imaging configurations optimized for human users typically involve digital processing of the captured images, which may lead to failures in re-identifying objects across different image frames. Summary of the Invention
[0006] An object of the present disclosure is to provide a system and method for improving the re-identification of objects in image frames of at least one video sequence.
[0007] According to a first embodiment, the present disclosure relates to a computer-implemented method for determining the appearance similarity of objects in image frames of at least one video sequence, comprising the following steps:
[0008] Process the first raw image data of the first image frame in a first image processing pipeline using a first image processing setting, thereby obtaining a processed first image;
[0009] Detect a first object in the processed first image frame and extract a first image region including the first object;
[0010] Extract one or more first feature vectors describing the appearance of the first object from the first image region;
[0011] Process the second raw image data of the second image frame in a second image processing pipeline using a second image processing setting, thereby obtaining a processed second image;
[0012] Detect a second object in the processed second image frame and extract a second image region including the second object;
[0013] Extract one or more second feature vectors describing the appearance of the second object from the second image region;
[0014] Determine the similarity between the first object and the second object by comparing one or more first feature vectors with one or more second feature vectors;
[0015] If one or more first feature vectors and one or more second feature vectors differ by more than a first threshold, and the first image processing setting and the second image processing setting differ by more than a second threshold, then reprocess the first raw image data corresponding to the first image region using the second image processing setting; extract one or more updated first feature vectors from the first image region;
[0016] Compare one or more updated first feature vectors with one or more second feature vectors.
[0017] In modern camera systems, there is usually an image processing pipeline in which raw image data is processed. For example, the image processing pipeline can reduce noise such as artifacts and image sensor noise, and adjust contrast, brightness, color, and other image attributes. The image processing pipeline can operate using a set of image processing settings, which can be affected by or depend on environmental conditions such as lighting or weather conditions, noise, etc. It has been found that changes in the image processing settings between frames can cause differences in the feature vectors extracted from the processed images, which in turn leads to failures in re-identifying objects.
[0018] It is possible to obtain one or more updated first feature vectors that can be re - compared with one or more second feature vectors by re - processing the first raw image data with the second image - processing setting if one or more of the second feature vectors differ by more than a first threshold and the first image - processing setting differs from the second image - processing setting by more than a second threshold. The updated first feature vectors (i.e., the feature vectors that have been re - processed with the image - processing setting of another image frame) can be used in the second comparison to determine whether the difference between the first and second feature vectors is caused by an actual feature - vector difference or by the fact that different image - processing settings are used for the two image frames.
[0019] The step of comparing one or more updated first feature vectors with one or more second feature vectors can be used to update the similarity determination of the first and second objects. An embodiment of a computer - implemented method for determining the appearance similarity of objects in image frames of at least one video sequence currently disclosed includes the step of updating the similarity determination of the first and second objects by comparing one or more updated first feature vectors with one or more second feature vectors.
[0020] To be able to re - process the first raw image data corresponding to the first image region with the second image - processing setting, it may be necessary to temporarily store the first raw image data, preferably in a storage medium such as random - access memory, at least until one or more of the first feature vectors and one or more of the second feature vectors have been compared. In the case where the first image - processing setting differs from the second image - processing setting by more than the second threshold, the first raw image data must be stored until the first raw image data has been re - processed. If the criteria for re - processing are not met, the first raw image data can be discarded; otherwise, the first raw image data can be discarded after re - processing. It is possible to temporarily store the first raw image data only for the first image region to avoid unnecessary memory usage. In the same way, it may be necessary to perform re - processing only on the first image region.
[0021] The computer - implemented method currently disclosed for determining the appearance similarity of objects in image frames of at least one video sequence can be executed repeatedly (e.g., at fixed time intervals, or for every nth image frame in the video sequence) to track objects. As those skilled in the art will recognize, the first / second raw image data, image frames, image settings, processed images, feature vectors, etc. can be extended to third, fourth, fifth raw image data, image frames, image settings, processed images, feature vectors, etc. with corresponding comparisons.
[0022] The present disclosure further relates to an image - processing system, which includes:
[0023] A processing circuit, configured to:
[0024] Apply a first image processing setting to process first raw image data of a first image frame in a first image processing pipeline, thereby obtaining a processed first image;
[0025] Detect a first object in the processed first image frame, and extract a first image region including the first object;
[0026] Extract one or more first feature vectors describing the appearance of the first object from the first image region;
[0027] Apply a second image processing setting to process second raw image data of a second image frame in a second image processing pipeline, thereby obtaining a processed second image;
[0028] Detect a second object in the processed second image frame, and extract a second image region including the second object;
[0029] Extract one or more second feature vectors describing the appearance of the second object from the second image region;
[0030] Determine the similarity between the first object and the second object by comparing one or more first feature vectors with one or more second feature vectors;
[0031] If one or more first feature vectors and one or more second feature vectors differ by more than a first threshold, and the first image processing setting and the second image processing setting differ by more than a second threshold, then apply the second image processing setting to re-process the first raw image data corresponding to the first image region; extract one or more updated first feature vectors;
[0032] Compare one or more updated first feature vectors with one or more second feature vectors.
[0033] The system may further include one or more image sensors, preferably in one or more cameras, for capturing image frames processed by the processing circuitry. The system may further include one or more displays for displaying the processed image frames, for example, in the form of one or more video sequences. Detected objects can be marked on the display. Alternatively or in combination, the processing circuitry may be configured to generate or extract metadata in which objects are marked according to, for example, the attributes of the objects, object identifiers, or the object classes to which they belong. If a method for determining the appearance similarity of objects determines that a first object in a first processed image frame is the same as a second object in a second processed image frame, or at least the likelihood is greater than a predefined threshold, the user can track the object with the appropriate label. Alternatively or in combination, the extracted information can be used for other applications, such as for analysis purposes.
[0034] Those skilled in the art will recognize that the currently disclosed methods for improving the re-identification of objects in the image frames of at least one video sequence can be performed using any embodiment of the currently disclosed image processing system, and vice versa. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In the following, various embodiments are described with reference to the accompanying drawings. The drawings are examples of embodiments and are intended to illustrate some of the features of the currently disclosed methods and systems for determining the appearance similarity of objects in the image frames of at least one video sequence.
[0036] Figure 1 An example of a flowchart showing the currently disclosed method for determining the appearance similarity of objects in the image frames of at least one video sequence is shown.
[0037] Figure 2 An illustration showing two objects in two different image frames is shown.
[0038] Figure 3 An example of an embodiment of the currently disclosed method for determining the appearance similarity of objects in the image frames of at least one video sequence is shown.
[0039] Figures 4A to 4B A possible configuration of the currently disclosed image processing system is shown. DETAILED DESCRIPTION
[0040] The present disclosure relates to a computer-implemented method for determining the appearance similarity of objects in the image frames of at least one video sequence. The present disclosure further relates to a method for tracking one or more objects in at least one video sequence using the currently disclosed method for determining the appearance similarity of objects in the image frames of at least one video sequence.
[0041] The method includes the following steps:
[0042] Process the first raw image data of the first image frame in a first image processing pipeline using a first image processing setting, thereby obtaining a processed first image;
[0043] Detect a first object in the processed first image frame and extract a first image region including the first object;
[0044] Extract one or more first feature vectors describing the appearance of the first object from the first image region.
[0045] The method further includes the following steps:
[0046] Process the second raw image data of the second image frame in a second image processing pipeline using a second image processing setting, thereby obtaining a processed second image;
[0047] Detect a second object in the processed second image frame and extract a second image region including the second object;
[0048] Extract one or more second feature vectors describing the appearance of the second object from the second image region.
[0049] The method may further include the following steps: the step of displaying the first object and / or the second object, preferably adding a label to indicate the identity of the tracked object.
[0050] Furthermore, if the currently disclosed method for determining the appearance similarity of objects determines that the similarity between objects in an image frame is higher than a similarity threshold, the first object may be associated with the second object. This may include, for example, providing an indication that the first object and the second object are considered to be the same object. This may include, for example, the step of outputting metadata indicating the association between the first object and the second object in the form of a common label.
[0051] Figure 1 An example showing an embodiment of the currently disclosed method 100 for determining the appearance similarity of objects in the image frames of at least one video sequence. In Figure 1 a specific example, a computer-implemented method 100 for determining the appearance similarity of objects in the image frames of at least one video sequence includes the following steps:
[0052] Process the first raw image data of the first image frame in a first image processing pipeline using a first image processing setting, thereby obtaining a processed first image (101);
[0053] Detect a first object in the processed first image frame and extract a first image region including the first object (102);
[0054] Extract one or more first feature vectors (103) that describe the appearance of the first object from the first image region;
[0055] Apply a second image processing setting to process the second raw image data of the second image frame in a second image processing pipeline, thereby obtaining a processed second image (104);
[0056] Detect a second object in the processed second image frame, and extract a second image region that includes the second object (105);
[0057] Extract one or more second feature vectors (106) that describe the appearance of the second object from the second image region;
[0058] Determine the similarity between the first object and the second object by comparing one or more first feature vectors with one or more second feature vectors (107);
[0059] If one or more first feature vectors and one or more second feature vectors differ by more than a first threshold, and the first image processing setting and the second image processing setting differ by more than a second threshold, then apply the second image processing setting to reprocess the first raw image data corresponding to the first image region; extract one or more updated first feature vectors from the first image region (108);
[0060] Compare one or more updated first feature vectors with one or more second feature vectors (109).
[0061] The steps of detecting a first object in a first processed image frame and detecting a second object in a second processed image frame can be accomplished in several ways. The steps can include, for example, using a convolutional neural network and can detect the objects substantially in real time. An example of a method for object detection is the method described in “You Only Look Once: Unified, Real-Time Object Detection” by J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, pages 779 - 788 of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, Nevada, USA, 2016 (doi: 10.1109 / CVPR.2016.91) (hereinafter referred to as YOLO). The YOLO method divides the input image into a grid of cells. For each cell, a single convolutional network is used to predict bounding boxes. For each bounding box, the likelihood that the detected object belongs to different predefined classes is calculated. Then, low-confidence bounding boxes can be filtered out. As will be recognized by those skilled in the art, the presently disclosed computer-implemented method for determining the appearance similarity of objects can use the YOLO method or other similar methods capable of detecting objects in an image.
[0062] As those skilled in the art will recognize, the steps of extracting one or more first feature vectors describing the appearance of a first object and / or extracting one or more second feature vectors describing the appearance of a second object can also be implemented in several ways. For example, the steps of extracting one or more first feature vectors and / or extracting one or more second feature vectors include applying a machine learning model such as a neural network. For this purpose, the method can use a convolutional neural network or the like. Preferably, the convolutional neural network or the like will be pre-trained to extract feature vectors. During the training of such a network, for example, it may be possible to use triplets of images. The triplets can include an anchor, a positive, and a negative. The triplets can include roughly aligned matching / mismatching objects. A loss function for a function that quantifies the difference between a predicted output and a ground truth can involve known techniques such as triplet loss and softmax loss. An example of a convolutional neural network for extracting feature vectors and its training is described in Schroff F., Kalenichenko D., and Philbin J.'s "Facenet: A Unified Embedding for Face Recognition and Clustering" (2015), Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 815 - 823, June 7 - 12, 2015, Boston (hereinafter referred to as FaceNet). Accordingly, the machine learning model can have been trained to compute feature vectors that are closer to each other for the same object and farther from each other for different objects. That is, the distance between feature vectors of the same object measured by a suitable distance metric (e.g., L2 norm) is shorter than the distance between feature vectors of different objects. Thus, the distance between feature vectors can be used to measure the similarity of the feature vectors and thereby the similarity of the objects, where the similarity increases as the distance between the feature vectors decreases.
[0063] The first image frame and the second image frame can be image frames of an image sensor of a camera at different time points, where the first image frame and the second image frame are image frames from a video sequence. The first image frame can be captured, for example, before the second image frame, or the second image frame can be captured before the first image frame. It is also possible that the first image frame and the second image frame are image frames from multiple image sensors. The multiple image sensors can be multiple image sensors depicting different views of a scene. For such an embodiment, the first image frame and the second image frame can still be image frames at different time points. Alternatively, the first image frame and the second image frame are image frames captured simultaneously. Some cameras have multiple image sensors for capturing a wider or even 360-degree panorama. The first image frame and the second image frame can be image frames from two of these image sensors, and as those skilled in the art will understand, it can be further extended to any number of image sensors and image frames. The multiple image sensors can also be sensors of different cameras. In an embodiment where the first image frame and the second image frame are image frames from multiple image sensors, the first image frame and the second image frame can be image frames from different video sequences.
[0064] The method can further include the step of determining the similarity between the first object and the second object by comparing one or more first feature vectors with one or more second feature vectors. In this regard, the feature vectors can be regarded as digital representations of the visual features of the objects extracted from the processed first image frame and the second image frame respectively. The term "feature vector" is commonly used in machine learning and will thus be known to those skilled in the art. One or more first feature vectors and one or more second feature vectors can also be referred to as re-identification vectors. Comparing one or more first feature vectors with one or more second feature vectors can include evaluating the similarity or dissimilarity between one or more first feature vectors and one or more second feature vectors. Many possible similarity metrics can be used for this purpose. The choice of similarity metric can depend on factors such as the content of the scene, the attributes of the objects, and the parameters and requirements related to re-identification. Possible similarity metrics include, but are not limited to, Euclidean distance, Manhattan distance, Jaccard similarity, Hamming distance, etc. It can be noted that the step of comparing one or more updated first feature vectors with one or more second feature vectors can perform the same comparison, i.e., the second comparison can apply, for example, the above similarity metrics. Accordingly, the step of comparing one or more first feature vectors with one or more second feature vectors and / or the step of comparing one or more updated first feature vectors with one or more second feature vectors can include using a similarity metric to calculate the similarity between one or more first feature vectors and one or more second feature vectors, and thereby also calculate the similarity between the first object and the second object.
[0065] Additional options when comparing one or more first feature vectors with one or more second feature vectors include a first step of determining to which object classes the first and second objects belong. This is a task that machine learning algorithms can generally perform with high accuracy. The step of determining to which object classes the first and second objects belong can actually be performed as part of the steps of detecting the first and second objects in the processed first and second image frames, respectively. Well-known object detectors such as YOLO and its further developments are generally able to determine the location and object class of the detected objects. The currently disclosed computer-implemented method for determining the appearance similarity of objects is not limited to any specific object detection algorithm or architecture. In some cases, the object detector only includes object localization and detection. For such applications, it is possible to use an additional classifier to determine to which object classes the first and second objects belong. Examples of architectures capable of performing object classification include AlexNet, VGG16, GoogleNets, EfficientNet, and Regnet. According to one embodiment, the method may include the steps of: determining whether the first and second objects belong to the same object class. If the first and second objects do not belong to the same class, it can be assumed that they are not the same object. Then, there is no need to further compare the feature vectors.
[0066] Although the object can be any object, the currently disclosed method for determining the appearance similarity of objects may be particularly useful for humans. In one embodiment, the first and second objects are humans. In many surveillance scenarios, being able to track individuals in one or more video sequences can be useful. Even if an individual turns around, changes body position, etc. between image frames, modern machine learning-based re-identification methods have the ability to re-identify and track the individual. If there is suspicion about whether the objects are the same, the currently disclosed method for determining the appearance similarity of objects can further improve this re-identification by ensuring that the comparison of the feature vectors is done using the same image processing settings. The currently disclosed method for determining the appearance similarity of objects can also be used in applications where the object is a vehicle. Accordingly, in one embodiment, the first and second objects are vehicles. Other applications are possible. For example, it is possible to re-identify and track animals, items in warehouses and stores, and transported goods.
[0067] If one or more first feature vectors and one or more second feature vectors differ by more than a first threshold, and additionally, a first image processing setting and a second image processing setting differ by more than a second threshold, then reprocess the first raw data. The first image processing setting and the second image processing setting differing by more than a second threshold can have the meaning that one or more parameters that are part of the processing setting (e.g., gain and / or exposure compensation and / or white balance and / or color adjustment such as local tone mapping) can be adjusted in a way that may target the quantization difference. As an example, "gain" in the context of processing a raw image can refer to the adjustment of the intensity or brightness of the image. Gain can be expressed in decibels (dB) that can be compared from one image processing setting to another. Other similar settings including those listed above also have similar levels. The reprocessing can be implemented on the first raw image data corresponding to the first image region. To reduce the computation and the data stored temporarily, it may only be necessary to reprocess the portion that covers the first image region. In this case, only the raw image data corresponding to the portion that covers the first image region needs to be temporarily stored for reprocessing. Perform the reprocessing on the first raw image data using the second image processing setting. The reprocessing generates a new processed image that can be referred to as an updated processed first image. Optionally, it may now also be possible to re-detect an object such as a first object in the updated processed first image and / or extract one or more updated first feature vectors that describe the appearance of the first object from the first image region. Then, the one or more updated first feature vectors can be re-compared with the one or more second feature vectors.
[0068] As will be commonly understood by those skilled in the art, the use of first / second raw image data, image frames, image settings, processed images, feature vectors, etc. does not necessarily have the meaning that the first image frame was captured before the second image frame. If, for example, the first image frame was captured before the second image frame, then it may be possible to apply the second image processing setting and reprocess the first raw image data corresponding to the first image region and extract one or more updated first feature vectors, but it will also be possible to apply the first image processing setting to reprocess the second raw image data corresponding to the second image region and extract one or more updated second feature vectors, and then the one or more updated second feature vectors can be compared with the one or more first feature vectors. Both variants can be useful. For this purpose, the first / second raw image data, image frames, image settings, processed images, feature vectors, etc. can refer to images from different time points, where "first" is not necessarily captured before "second".
[0069] The first image region including the first object and the second image region including the second object may respectively refer to regions or sub-regions of the first image frame and the second image frame. Extracting regions or sub-regions of an image frame is sometimes also referred to as cropping. In the present disclosure, the first image region and the second image region may refer to regions encapsulating the first object / second object, where the region is smaller than the region of the entire first image frame / second image frame. The first image region and the second image region are generally but not necessarily rectangular or square.
[0070] Figure 2 An example showing an embodiment of a method for determining the appearance similarity of objects in an image frame of at least one video sequence according to the present disclosure is presented. In this example, the left image frame is the first image frame 202. The right image frame is the second image frame 201. As can be seen, the first image frame 202 includes two first objects 205 and 209. The second image frame 201 includes two second objects 203 and 207. The first image region 206 in the first image frame 202 includes the first object 205. Another first image region 210 includes another first object 209. The second image region 204 in the second image frame 201 includes the second object 203. Another second image region 208 includes another second object 207. In the process of re-identifying objects from one image frame to another, it can be seen from this example that the second object 203 in the second image frame 201 may be the first object 205 or another first object 209 that has moved from the first image frame 202 to a new position in the second image frame 201. The computer-implemented method for determining the appearance similarity of objects in an image frame according to the present disclosure can be used to evaluate the similarity between the second object 203 and the first object 205 and / or another first object 209 to re-identify one of the first objects 205 and 209 as being the same as the second object 203.
[0071] In one embodiment of the currently disclosed method for determining the appearance similarity of objects, the step of comparing one or more first feature vectors with one or more second feature vectors and / or the step of comparing one or more updated first feature vectors with one or more second feature vectors includes applying a machine learning model such as a neural network that has been trained to compare the appearance of two objects based on the feature vectors describing the appearance of the two objects. Training a machine learning model such as a neural network typically may include collecting a training data set including annotations or other information indicating which objects are the same or different in different frames. As previously mentioned, the machine learning model can be trained, for example, by using triplet training on images of objects known to be the same or different to calculate feature vectors that are closer to each other for the same objects and farther from each other for different objects. When one or more first feature vectors and one or more second feature vectors have been extracted, the comparison of the feature vectors can be completed using a similarity / dissimilarity metric. As mentioned above, an example of a similarity metric is the Euclidean distance. In an example of using the FaceNet system to measure similarity, the square of the L2 distance is used to determine object similarity. The feature vectors can be provided in any suitable format such as an array, a database entry, a text file, a binary format, etc. Training may involve other well-known steps such as data augmentation, transfer learning, validation, etc.
[0072] Figure 3 An example showing an embodiment of the currently disclosed method for determining the appearance similarity of objects in the image frames of at least one video sequence is presented. In this example, for a first object 205 in a first image region 206 of a first image frame 202, a plurality of first feature vectors 212 describing the appearance of the first object 205 are extracted. For a second image frame 201, a plurality of second feature vectors 211 describing the appearance of a second object 203 are extracted. The next step is to determine the similarity between the first object 205 and the second object 203 by comparing one or more first feature vectors 212 with one or more second feature vectors 211. If the first feature vectors 212 and the second feature vectors 211 differ by more than a first threshold and, at the same time, the first image processing setting differs from the second image processing setting by more than a second threshold, then the second image processing setting is applied and the first original image data corresponding to the first image region 206 is reprocessed, and vice versa. Based on this, updated first feature vectors can be obtained, and the updated first feature vectors can be re-compared with one or more second feature vectors.
[0073] As described above, in one embodiment of the currently disclosed method for determining the appearance similarity of objects, the first raw image data corresponding to at least the first image region is temporarily stored until one or more first feature vectors and one or more second feature vectors have been compared, and in the case where the first image processing setting differs from the second image processing setting by more than a second threshold, the first raw image data corresponding to at least the first image region is temporarily stored until the first raw image data has been reprocessed. The first raw image data and / or the second raw image data may be temporarily stored until one or more first feature vectors have been compared with one or more second feature vectors and / or until the first raw image data has been reprocessed. After the comparison, the raw image data may be discarded. In one embodiment, after one or more first feature vectors have been compared with one or more second feature vectors and / or the first raw image data has been reprocessed, the first raw image data and / or the second raw image data is discarded.
[0074] The present disclosure further relates to a computer program having instructions that, when executed by a computing device or a computing system, cause the computing device or the computing system to implement any embodiment of the currently disclosed method for determining the appearance similarity of objects. The computer program may be stored on any suitable type of storage medium such as a non-transitory storage medium.
[0075] As described above, the present disclosure further relates to an image processing system that includes processing circuitry configured to execute any embodiment of the currently disclosed method for determining the appearance similarity of objects. The processing circuitry may include a single image processing pipeline, where the first image processing pipeline is the same image processing pipeline as the second image processing pipeline, or a separate image processing pipeline for processing the first raw image data of the first image frame and the second raw image data of the second image frame. The latter may be useful, for example, if there are multiple cameras for capturing different views of a scene.
[0076] The image processing system may include a machine learning model such as a neural network that has been trained to compare the appearance of two objects based on feature vectors describing the appearance of the two objects.
[0077] The system may (but does not necessarily have to) include a display for displaying the processed image frames of at least one video sequence. When tracking an object, it may be useful to display the object with, for example, a label or a bounding box. However, in other applications, the re-identified object is not necessarily displayed but is used for additional applications including, for example, further analysis.
[0078] The system may further include peripheral components such as one or more memory units that may be used to store instructions executable by the processing circuitry. The system may further include internal and external network interfaces, input and / or output ports, etc.
[0079] As will be understood by those skilled in the art, the processing circuitry may include a single processor or processing unit in a multi-core / multi-processor system. The processing circuitry may be connected to a data communication infrastructure.
[0080] The system may include one or more memory units such as random access memory (RAM) and / or read-only memory (ROM) or any suitable type of memory. The system may further include a communication interface that allows software and / or data to be transferred between the system and external devices. The software and / or data transferred via the communication interface may be in any suitable form of electrical, optical, or RF signal. The communication interface may include, for example, a cable or wireless interface.
[0081] The camera may be connected to the processing circuitry via one or more Ethernet cables that may use the power of Ethernet.
[0082] Figures 4A to 4B Two possible configurations of the presently disclosed image processing system 300 are shown. In Figure 4A the example, the image processing system 300 includes processing circuitry 301 and a camera 302 for capturing image frames including at least a first image frame and a second image frame of at least one video sequence. On the right side of the example, how an object changes between the two image frames can be seen. The camera 302 may include only one image sensor or multiple image sensors. In Figure 4B the example, the image processing system 300 includes processing circuitry 301 and two cameras 302 for capturing image frames including at least a first image frame and a second image frame of at least one video sequence. On the right side of the example, how an object changes between the two image frames can be seen. It can be noted that the processing circuitry may have a single image processing pipeline, but it is also possible that each camera has its own image processing pipeline.
Claims
1. A computer-implemented method for determining appearance similarity of objects in image frames of at least one video sequence, the method comprising: Applying a first image processing setting to process first raw image data of a first image frame in a first image processing pipeline to obtain a processed first image; detecting a first object in the processed first image frame, and extracting a first image region including the first object; extracting one or more first feature vectors describing the appearance of the first object from the first image region; Applying a second image processing setting to process second raw image data of a second image frame in a second image processing pipeline, thereby obtaining a processed second image; detecting a second object in the processed second image frame, and extracting a second image region including the second object; extracting one or more second feature vectors describing the appearance of the second object from the second image region; determining a similarity between the first object and the second object by comparing the one or more first feature vectors with the one or more second feature vectors; If the one or more first feature vectors and the one or more second feature vectors differ by more than a first threshold, and the first image processing setting differs from the second image processing setting by more than a second threshold, reprocessing the first original image data corresponding to the first image area by applying the second image processing setting; extracting one or more updated first feature vectors from the first image area; The one or more updated first feature vectors are compared to the one or more second feature vectors.
2. The computer-implemented method for determining appearance similarity of objects according to claim 1, wherein: Temporarily storing the first original image data corresponding to at least the first image area until the one or more first feature vectors and the one or more second feature vectors have been compared, and in the event that the first image processing setting differs from the second image processing setting by more than a second threshold, temporarily storing the first original image data corresponding to at least the first image area until the first original image data has been reprocessed.
3. The computer-implemented method for determining appearance similarity of objects according to claim 2, wherein: The first raw image data and / or the second raw image data are temporarily stored until the one or more first feature vectors have been compared with the one or more second feature vectors and / or until the first raw image data have been reprocessed.
4. The computer-implemented method for determining appearance similarity of objects according to claim 2, wherein: After the one or more first eigenvectors have been compared to the one or more second eigenvectors and / or the first raw image data have been reprocessed, the first raw image data and / or the second raw image data are discarded.
5. The computer-implemented method for determining appearance similarity of objects according to claim 1, wherein: The first image frame is captured before the second image frame, or the second image frame is captured before the first image frame.
6. The computer-implemented method for determining appearance similarity of objects according to claim 1, wherein: The step of comparing the one or more first feature vectors with the one or more second feature vectors includes determining whether the first object and the second object belong to the same object class.
7. The computer-implemented method for determining appearance similarity of objects according to claim 1, wherein: The one or more first feature vectors and the one or more second feature vectors are numerical representations of features extracted from the processed first image frame and the processed second image frame, respectively.
8. The computer-implemented method for determining appearance similarity of objects according to claim 1, wherein: The steps of extracting one or more first feature vectors and / or extracting one or more second feature vectors include applying a machine learning model such as a neural network.
9. The computer-implemented method for determining appearance similarity of objects according to claim 8, wherein: The machine learning model has been trained to compute feature vectors that are closer to each other for the same objects and farther away from each other for different objects.
10. The computer-implemented method for determining appearance similarity of objects according to claim 1, wherein: The step of comparing the one or more first feature vectors with the one or more second feature vectors and / or the step of comparing the one or more updated first feature vectors with the one or more second feature vectors comprises calculating a similarity between the first object and the second object using a similarity measure.
11. The computer-implemented method for determining appearance similarity of objects according to claim 1 , wherein: The first image processing settings and the second image processing settings include settings related to gain and / or exposure compensation and / or white balance and / or color adjustment such as local tone mapping.
12. The computer-implemented method for determining appearance similarity of objects according to claim 1, wherein: The first image frame and the second image frame are image frames from an image sensor of a camera at different time points, wherein the first image frame and the second image frame are image frames from a video sequence.
13. The computer-implemented method for determining appearance similarity of objects according to claim 1, wherein: The first image frame and the second image frame are image frames from a plurality of image sensors of a camera, wherein the first image frame and the second image frame are image frames from different video sequences.
14. A computer program having instructions which, when executed by a computing device or a computing system, cause the computing device or the computing system to implement the method for determining appearance similarity of objects according to claim 1.
15. An image processing system, comprising: The processing circuit is configured to: Applying a first image processing setting to process first raw image data of a first image frame in a first image processing pipeline to obtain a processed first image; detecting a first object in the processed first image frame, and extracting a first image region including the first object; extracting one or more first feature vectors describing the appearance of the first object from the first image region; Applying a second image processing setting to process second raw image data of a second image frame in a second image processing pipeline, thereby obtaining a processed second image; detecting a second object in the processed second image frame, and extracting a second image region including the second object; extracting one or more second feature vectors describing the appearance of the second object from the second image region; determining a similarity between the first object and the second object by comparing the one or more first feature vectors with the one or more second feature vectors; If the one or more first eigenvectors and the one or more second eigenvectors differ by more than a first threshold, and the first image processing setting differs from the second image processing setting by more than a second threshold, reprocessing the first original image data corresponding to the first image area by applying the second image processing setting; extracting one or more updated first eigenvectors; The one or more updated first feature vectors are compared to the one or more second feature vectors.