Method of object re-identification

By training multiple neural networks on different anatomical features and combining predefined conditions and quality assessment, the problem of neural networks being unable to re-identify objects after they are occluded or disappearing is solved, thus improving the accuracy and robustness of object recognition.

CN112784669BActive Publication Date: 2025-10-17AXIS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011144403.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-08
Filing Date
2020-10-23
Publication Date
2025-10-17
Estimated Expiration
2040-10-23

AI Technical Summary

Technical Problem

Existing neural networks struggle to effectively re-identify objects after they are occluded or disappear, especially when objects are occluded to varying degrees or disappear from the scene for short or long periods. This can lead to objects being identified as new and different objects, interfering with video analysis algorithms.

Method used

Multiple neural networks are employed, each trained on different groups of anatomical features. By receiving multiple images of an object, the most similar reference vector is determined and input into the corresponding neural network to identify the object. Predefined conditions and quality assessments are combined to ensure recognition accuracy.

Benefits of technology

It improves the accuracy and robustness of object re-identification, reduces identification errors caused by object occlusion or disappearance, and enhances the reliability of video analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112784669B_ABST
    Figure CN112784669B_ABST
Patent Text Reader

Abstract

A method of object re-identification is disclosed, in particular a method of object re-identification in images of objects. The method comprises providing a plurality of neural networks (27) for object re-identification, wherein each of the plurality of neural networks is trained on image data of a different group of anatomical features, each group being represented by a reference vector; receiving a plurality of images (4) of an object (38) and an input vector representing anatomical features depicted in all of the plurality of images (4); comparing the input vector to the reference vectors to determine a most similar reference vector according to a predefined condition; and inputting image data of the plurality of objects (38) to the neural network (#1) represented by the most similar reference vector to determine whether the plurality of objects (38) have the same identity.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to the field of object re-identification by means of neural networks. BACKGROUND

[0002] Object re-identification techniques have been widely studied and used, for example, to identify and track objects in related digital images.

[0003] It is well known that humans can easily identify and relate objects of the same identity in images, even when the objects become occluded to different extents or even disappear from a scene for a short or long time. The appearance of an object can also vary according to the viewing angle and over time. However, the re-identification of objects is challenging for computer vision systems, especially in scenes where the objects become occluded (i.e. not fully visible) or disappear completely from a scene and appear later in the same scene or in another scene.

[0004] For example, one challenge is to resume tracking of an object when the object has left a scene and entered the same scene or another scene monitored by another camera. If the tracking algorithm cannot resume tracking, the object will be identified as a new, different object, which can interfere with other algorithms for video analysis.

[0005] There are proposals to use neural networks to assist in re-identification. However, there is a need to provide an improved method and apparatus for re-identifying objects in images and videos. SUMMARY

[0006] The present invention aims to provide a method of re-identification by means of neural networks. As mentioned above, the use of neural networks for re-identification brings possible drawbacks. For example, a neural network trained on images of complete body structures can not be able to re-identify a person in an image frame in which only the upper body part of the body structure is visible. It has also been shown that neural networks struggle to successfully perform re-identification based on images showing different amounts of the object, for example images showing the upper body in some of the images and images showing the full body in some of the images. This can be the case, for example, when monitoring a scene in which people are entering the scene (showing the full body), sitting down (showing the upper body) and leaving the scene (again showing the full body, but possibly at a different angle).

[0007] Therefore, the inventors have realized that one drawback of object re-identification is the difficulty to re-identify an object from images showing different amounts of the object. This has been found to be a problem, for example, for human objects.

[0008] It is an object of the present invention to eliminate or at least reduce this and other drawbacks of the presently known methods of object re-identification for objects, in particular for human objects.

[0009] According to a first aspect, these and other objects are all or at least partially achieved by a method of object re-identification in images of objects, the method comprising:

[0010] • providing a plurality of neural networks for object re-identification, wherein each of the plurality of neural networks is trained on image data having a different set of anatomical features, each set being represented by a reference vector,

[0011] • receiving a plurality of images of an object and an input vector representing anatomical features depicted in all of the plurality of images,

[0012] • comparing the input vector with the reference vectors to determine the most similar reference vector according to a predefined condition,

[0013] • inputting image data of a plurality of objects to the neural network represented by the most similar reference vector to determine whether the plurality of objects have the same identity. The same identity means that the plurality of objects imaged in the plurality of images are in fact the same object imaged multiple times.

[0014] The present invention is based on the recognition that known neural networks trained on object re-identification can struggle to perform well when the input image data comprises objects that are visible to different degrees. In other words, re-identification often fails when the object of the input data is more or less occluded in the images of the input image data. The inventors have come up with a solution of training different neural networks on reference data that is uniform with respect to the amount of depicted objects. In other words, different neural networks have been trained on different sets of anatomical features for the type of object. Depending on the image data on which re-identification is to be performed based, the appropriate neural network is selected. In particular, the neural network trained on data having a set of anatomical features that satisfy a predefined condition is selected. The predefined condition is a similarity condition that defines the degree of similarity that the compared vectors have. Prior to selecting the neural network, an input vector of the image data is determined. The input vector represents the anatomical features depicted in all images of the image data. This input vector is compared to the reference vectors of the neural networks, wherein each reference vector represents the anatomical features of the reference data of its respective neural network. By adding this solution as a pre-step to inputting image data to a neural network for re-identification, the performance of the re-identification is improved without the need for complex algorithms for estimating non-depicted parts of the object, for example. It is relatively uncomplicated to implement the solution of the present invention by using known algorithms for determining the anatomical features depicted in all of the plurality of images and by referring to known neural network structures for re-identification.

[0015] The object is of a type that can be re-identified by image analysis. This means that individuals or groups of individuals of the object type can be separated from each other based on appearance. Each individual of the object type does not need to be uniquely identifiable relative to all other individuals of the object type. For the method of the invention to be beneficial, it is enough that there is a difference between some individuals or groups of individuals.

[0016] The object type can be a human being. In such embodiments, the method is directed to re-identification of human objects. Other non-limiting examples of object types are vehicles, animals, baggage objects such as suitcases, backpacks, handbags and other types of bags, and parcels including letters. The method can be extended to perform re-identification on larger objects such as buildings and landmarks, as long as they can be re-identified by image analysis as defined above.

[0017] An anatomical feature, in the context of the present application, refers to a distinct unique part of an object. For a human body, anatomical features include, for example, nose, eye, elbow, neck, knee, foot, shoulder, and hand. One part can have different appearances between different objects. For example, a foot with shoes or without shoes or a foot wearing different styles of shoes, although having different appearances, are still considered to be the same anatomical feature. For a vehicle, anatomical features include, for example, window frame, wheel, tail light, side mirror, and sunroof. Distinct parts mean that the anatomical features do not overlap with each other. For example, a human arm includes different distinct anatomical features such as shoulder, upper arm, elbow, forearm, wrist, and back of hand. An anatomical feature can be seen as corresponding to a distinct physical point on the object, where the anatomical feature is represented by the part of the object surrounding the respective point.

[0018] The input vector / reference vector refers to a representation vector representing the input values / reference values of the anatomical features. Depending on how the anatomical features are determined and thus represented, e.g. by key points, the input vector / reference vector can have different forms. Thus, the representation can differ between different implementations, which is a known fact that can be handled by the skilled person based on existing knowledge. As an example, the input vector / reference vector can have the form of a one-dimensional vector with numerical values. The input vector / reference vector can be a vector with binary values, where each position in the vector represents an anatomical feature. For example, a 1 in a particular position in the vector can indicate that the respective anatomical feature is detected / visible, while 0 can indicate that the respective anatomical feature is not detected / visible.

[0019] The input vector can be an edge vector (representing edges of the object), a contour vector (representing contours of the object), or a key point vector representing key points of a human object. It is well known that key points are commonly used for object detection and processing image data. Key points of an object can be found by using a neural network. The key points can represent anatomical features.

[0020] The edges or contours of the objects provide an alternative way of representing the objects in the image data. How to determine depicted edges or contours of objects in given image data is well known, for example by the known Sobel, Prewitt and Laplacian methods. Edges and contours can be determined by using neural networks designed and trained for this purpose. From the edges or contours, anatomical features can be determined.

[0021] The predefined condition can define that the reference vector that is equal to the input vector is determined as the most similar reference vector. In other words, in this embodiment the most similar reference vector is the reference vector that is equal to the input vector. The respective neural network associated with this reference vector should then be used for the re-identification. In this embodiment, the selected neural network is trained on an image that comprises the same anatomical features as comprised by all images in the input image data (i.e. in the plurality of images).

[0022] The predefined condition can define that from the reference vectors the reference vector is determined as the most similar reference vector that has the largest overlap with the input vector. The neural network corresponding to such a reference vector is trained on image data that has all anatomical features represented in the plurality of images. This embodiment can form a second choice of the previously disclosed embodiments. That is, the method can first try to find a reference vector that is equal to the input vector, and when this is not successful, select the reference vector that has the largest overlap with the training vector. Other conditions can also be included, for example the input vector needs to satisfy certain quality conditions that will be disclosed later.

[0023] If there is more than one reference vector that satisfies the similarity condition (is equal or has the same amount of overlap), the predefined condition can include further selection criteria. For example, some of the anatomical features represented by the input vector can have a greater influence on the re-identification than other anatomical features, and then the reference vector that represents one or more important anatomical features is selected before other reference vectors. Another example is to select the largest matching subset between the input vector and the reference vector among the reference vectors that satisfy the other criteria of the selection criteria.

[0024] The predefined condition can define to determine from the reference vectors the reference vector comprising the largest number of anatomical features defined by the priority list overlapping with the input vector. In other words, the input vector is compared to the reference vectors to find the reference vector having the largest overlap with the set of anatomical features included in the priority list. The priority list is predefined and can list anatomical features known to increase the chance of a successful re-identification. Such anatomical features can include eyes, nose, mouth, shoulder, etc. The priority list can differ between different applications and can be related to the configuration of the neural network or to feedback on the performance of the neural network. For example, if it is determined that the neural network performs particularly well on image data comprising a shoulder, this anatomical feature is added to the priority list. Thus dynamic updating of the priority list based on feedback can be implemented.

[0025] The method can further comprise:

[0026] • evaluating the input vector against a preset quality condition,

[0027] • when the preset quality condition is met, performing the step of comparing the input vector to the input image data, and

[0028] • when the preset quality condition is not met, discarding at least one image of the plurality of images, determining a new input vector based on the reduced plurality of images as the input vector, and iterating the method from the step of evaluating the input vector.

[0029] This embodiment adds a quality assurance to the method. Even with the suggested method, where a suitable neural network for re-identification is selected, a low quality of the input data can decrease the performance of the neural network. By ensuring that the input data has a certain quality, a minimum level of performance can be maintained. The preset quality condition can be for example a minimum vector size.

[0030] The evaluation of the input vector against the preset quality condition can comprise the action of comparing the input vector to a predefined list of anatomical features at least one of which should be represented in the input vector.

[0031] If the condition is not met, the method can comprise the further action of discarding one or more of the plurality of images and iterating the method based on the reduced plurality of images. The images that are discarded can be selected based on their content. For example, images that do not comprise any of the anatomical features of the predefined list can be discarded. This discarding step can be performed prior to evaluating the input vector to make the method faster.

[0032] The plurality of images can be captured by one camera at a plurality of time points. The plurality of images thus forms an image sequence depicting a scene. In another embodiment, the plurality of images can be captured by a plurality of cameras covering the same scene from different angles. The plurality of images thus forms a plurality of image sequences. In yet another embodiment, the plurality of images can be captured by a plurality of cameras depicting different scenes, which also results in a plurality of image sequences.

[0033] It can be interesting to perform re-identification in each of these scenarios, however the purpose and application of re-identification can differ. Re-identification can for example assist an object tracking algorithm that is more generally applied to monitor a single scene rather than different scenes. In such an embodiment, the purpose of re-identification can be to ease re-tracking of a person after the person has been occluded.

[0034] In another case, cameras monitor the same scene from different angles. The plurality of images can be taken at the same time point. The purpose of re-identification can be to connect images comprising the same object but acquired by different cameras.

[0035] In the case with different scenes, each scene is monitored by a camera, the plurality of images can be collected from different cameras. In this case, the purpose of re-identification can be long term tracking, where a person leaves one scene and can appear in another scene after a few minutes, hours or even days. The scenes can for example be different parts of a city, and the purpose of re-identification can be to track a wanted person or vehicle.

[0036] The image data of the input plurality of images can comprise inputting image data representing only anatomical features depicted in all of the plurality of images. In this embodiment, the method can comprise the action of filtering the image data of the plurality of images based on anatomical features depicted in all of the plurality of images, before the step of inputting the image data to the selected neural network.

[0037] As part of the step of receiving the plurality of images, the method can further comprise:

[0038] • acquiring the plurality of images by one or more cameras,

[0039] • determining anatomical features depicted in all of the plurality of images, and

[0040] • determining an input vector representing the determined anatomical features.

[0041] In other words, the method can comprise an initial process of forming a plurality of images. According to this embodiment, the plurality of images can be prepared by another processor than the processor executing the main part of the method, i.e. the comparison of the input vector with the reference vectors to determine the neural network. Alternatively, the preparation can be carried out within the same processing unit. The result of the initial process, being the input vector and the plurality of images, can be transmitted within or to the processing unit to be executed the subsequent method steps.

[0042] The step of receiving a plurality of images in the method can comprise:

[0043] • capturing images by one or more cameras, and

[0044] • selecting different images to form the plurality of images based on predetermined frame distance, time gap, image sharpness, pose of depicted subject, resolution, aspect ratio of region and plane rotation.

[0045] In other words, as an initial step of the main method of determining a suitable neural network, images of suitable candidates for re-identification can be filtered out. The purpose of the filtering can be to select images that can have the same subject, and / or images on which the method can be well performed.

[0046] According to a second aspect, the above mentioned objects and other objects are fully or at least partially achieved by a non-transitory computer-readable recording medium having recorded thereon computer readable program code which, when executed on a device having processing capabilities, is configured to perform any of the above disclosed methods.

[0047] According to a second aspect, the above mentioned objects and other objects are fully or at least partially achieved by a controller for controlling a video processing unit for subject re-identification. The controller can have access to a plurality of neural networks for subject re-identification, wherein each of the plurality of neural networks is trained on image data of different groups of anatomical features, each group being represented by a reference vector. The controller comprises:

[0048] • a receiver configured to receive a plurality of images of a human subject and an input vector representing anatomical features depicted in all of the plurality of images,

[0049] • a comparison component adapted to compare the input vector with the reference vectors to determine the most similar reference vector according to predefined conditions,

[0050] • a determination component configured to input image data of the plurality of subjects to the neural network represented by the most similar reference vector to determine whether the plurality of human subjects have the same identity, and

[0051] • a control component configured to control whether the video processing unit considers the plurality of objects as having the same identity.

[0052] The image processing unit of the third aspect can typically be implemented in the same way as the method of the first aspect, with the attendant advantages.

[0053] Further areas of applicability of the present application will become apparent from the detailed description given below. It should be understood, however, that the description and specific examples, while indicating the preferred embodiment of the application, are given by way of illustration only, since various changes and modifications within the scope of the application will become apparent to those skilled in the art from this detailed description.

[0054] It should be understood, therefore, that the present application is not limited to the particular components described or the steps of the methods described, as such equipment and methods can vary. It should be further understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It must be noted that, as used in the specification and the appended claims, the articles "a," "an," and "the" are intended to mean one or more of the elements in the context of each use thereof. Thus, for example, reference to "a subject" or "the subject" can include a plurality of subjects. Also, the words "comprise," "comprises," and "comprising" are not intended to exclude the presence of elements other than the elements expressly listed. The term "consisting essentially of to indicate the presence of other elements in addition to the elements expressly stated or illustrated. BRIEF DESCRIPTION OF DRAWINGS

[0055] The present application will now be described in more detail, by way of example, and with reference to the drawings, in which:

[0056] Figure 1 Flowcharts illustrating different embodiments of a method of object re-identification are shown.

[0057] Figure 2 An overall overview of the method is provided.

[0058] Figure 3 An image sequence is illustrated.

[0059] Figure 4 A plurality of images selected from Figure 3 an image sequence are illustrated.

[0060] Figure 5 A pair of images captured from different angles of a scene are illustrated.

[0061] Figure 6 A plurality of images selected from different image sequences are illustrated. DETAILED DESCRIPTION

[0062] Reference is first made to Figure 1 and Figure 2Summary of the disclosed method. Reference will be made to Figure 1 selected steps of which will be disclosed later. The purpose of the method is to re-identify an object based on images captured by one or more cameras. As discussed earlier, the purpose of the re-identification can vary from application to application.

[0063] Correspondingly, the method comprises a step S102 of capturing an image 22 by at least one camera 20. The camera 20 monitors a scene 21. In this embodiment, an object in the form of a human being appears in the scene and is imaged by the camera 20. The image 22 is processed by a processing unit 23, which can be located in the camera 20 or as a separate unit in wired or wireless connection with the camera 20. The processing unit 23 detects S104 the object in the image 22 by means of an object detector 24. This can be performed by a well-known object detection algorithm. The algorithm can be configured to detect a specific type of object, such as a human object.

[0064] Then, a step S105 of selecting a plurality of images from the image 22 can be performed. Alternatively, the step S105 can be performed before the step S104 of detecting the object in the image 22. Details of the selection step S105 will be disclosed later.

[0065] Based on the plurality of images, an anatomical feature is determined by the processing unit 23, more precisely by a feature extractor 26. The determination of the anatomical feature can be performed by executing a well-known image analysis algorithm. For example, a system known as “OpenPose” (disclosed in “OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields” by Cao et al.) can be used. OpenPose is a real-time system capable of detecting body and hand keypoints on a single image.

[0066] Depending on the image analysis technique of the application, the determined anatomical feature can be represented differently. Examples of representations are by keypoints (e.g. in the form of a vector of keypoints), by edges (e.g. in the form of a vector of edges) or by contours (e.g. in the form of a vector of contours).

[0067] Next, the processing unit 23 analyses the plurality of images and / or the representation of the determined anatomical feature and determines S108 an input vector representing the anatomical feature that is represented in all of the plurality of images.

[0068] The optional steps of evaluating S109 the input vector and discarding S111 one or more images will be disclosed in detail later.

[0069] At the core of the inventive concept, after the input vector has been determined, the input vector is compared S112 with reference vectors representing training data on which the set 29 of neural networks #1, #2, #4, #3, #5 has been trained. The neural networks are provided S110 to the processing unit 23, meaning that they can be used by the processing unit 23. They can be in the form of separate neural networks, or in the form of neural networks included in a single neural network structure 27, where the different neural networks are formed by different connections or paths in the neural network structure. The neural networks have been trained on different training data (represented by different reference vectors). The reference vectors are provided in a format that enables comparison with the input vector. For example, both the input vector and the reference vectors can be in the form of keypoint vectors. Or, the input vector can be a keypoint vector, while the reference vectors can be object landmark vectors or stick figures, for which conversion to keypoint vector format can be performed in a straightforward manner.

[0070] The comparison S112 is performed by a comparator 28 of the processing unit 23. The purpose of the comparison S112 is to find the reference vector that is most similar to the input vector. Similarity is defined by a pre-defined condition. Examples of such conditions will be disclosed in detail later. Based on the result of the comparison, a neural network (in the illustrated example #1) is selected. Thus, the neural network that has been trained on image data with anatomical features that are most similar to the anatomical features represented by the input vector is selected. All or a selected portion of the image data from the plurality of images is input S116 to the selected neural network (#1).

[0071] The result from the selected neural network is received S118 by the processing unit 23. In other embodiments, the result of the re-identification can be sent to other units, such as a separate control unit. The processing unit 23 can alternatively form part of a control unit or controller (not illustrated).

[0072] In this example, however, the processing unit 23 receives S118 the result from the neural network (#1). In essence, the result provides information about whether the object of the plurality of images has the same identity. The processing unit 23 uses this information to control the camera 20. For example, the information can be used by the camera 20 for continuing to track the object after the object has been occluded.

[0073] In one embodiment, the method further comprises determining a pose for each detected object. For example, a pose can be determined for a human object based on anatomical features such as key points. The determined pose can be included in the input vector. In such an embodiment, the reference vector further comprises pose data corresponding to the poses of objects in the image data on which the network has been trained. This feature can further assist in selecting a neural network for re-identification that is suitable for the current input vector.

[0074] The functions of the processing unit 23 can be implemented as hardware, software, or a combination thereof.

[0075] In a hardware implementation, the components of the processing unit (e.g. the object detector 24, the feature extractor 26 and the comparator 28) can correspond to dedicated and specifically designed circuits to provide the functionality of the components. The circuits can be in the form of one or more integrated circuits, such as one or more application-specific integrated circuits or one or more field-programmable gate arrays.

[0076] In a software implementation, the circuits can alternatively be in the form of a processor, such as a microprocessor, associated with computer code instructions stored on a (non-transitory) computer-readable medium, such as a non-volatile memory, causing the processing unit 23 to perform (part of) any of the methods disclosed herein. Examples of non-volatile memory include read-only memory, flash memory, ferroelectric RAM, magnetic computer storage devices, and optical discs, etc. In the case of software, the components of the processing unit 23 can thus each correspond to a portion of the computer code instructions stored on the computer-readable medium that, when executed by the processor, cause the processing unit 23 to perform the functionality of the components.

[0077] It will be appreciated that there can also be a combination of hardware and software implementations, meaning that the functionality of some of the components in the processing unit 23 is implemented in hardware, while other components are implemented in software.

[0078] In more detail, reference will now be made to further embodiments of the processing unit 23. Figure 3 and Figure 4 The method is disclosed. Figure 3The diagram shows a sequence of images acquired by a single surveillance camera monitoring a scene. The image sequence includes digital images 31-36 and is organized in chronological order. The image sequence images an event process in which a person 38 is about to cross a pedestrian crossing on a road 39 when a truck 37 neglects to give way, which (needless to say) irritates the person 38 who must suddenly move to the side before continuing to cross the road 39. When the truck 37 drives past the person 38, the latter is obscured by the truck 37 from the camera's perspective. After the occlusion, a tracking algorithm attempting to track the person 38 will likely be unable to continue tracking the person 38. Instead, after the occlusion, the person 38 will be detected as a new object with a new identity. Re-identification can help to correct this defect.

[0079] According to this method, from Figure 3 Select S105 from the image sequence Figure 4 , namely, images 31, 32, and 34. These images 31, 32, 34 may be selected based on different selection criteria. For example, images depicting one or more objects may be selected. Other non-limiting examples of selection criteria for selecting which images in a set of images will form the plurality of images include:

[0080] A predetermined frame interval, such as every 90th frame.

[0081] Time intervals, such as every 5th second.

[0082] Image sharpness, which can be determined by determining the sharpness for each image and selecting the image with the best sharpness. Sharpness can be determined for the entire image or for a selected area of ​​the image (eg where an object is or may be located).

[0083] The pose of the detected object, which can be determined by looking at the key points, edges, or outlines of the detected object. Images with objects having a specific pose or having similar poses can be selected.

[0084] Resolution: You can determine the resolution for the entire image or a selected area. The image with the best resolution is selected.

[0085] The aspect ratio of the object area, where the area may correspond to a bounding box. The aspect ratio provides information about the size of the object. Different aspect ratios may be suitable for different applications.

[0086] Next, object detection is performed on the plurality of images 4. In this example, one object in each image 31, 32, 34 is detected. The purpose of this method is to determine whether the objects have the same identity. A set of common anatomical features of the detected objects in the plurality of images is determined, i.e. anatomical features that are depicted in all of the plurality of images 4. This set of common anatomical features can be determined by determining key points and is represented by an input vector. As disclosed above, the input vector is then compared S112 to reference vectors that are associated with available neural networks that can be used for re-identification of the detected objects in the plurality of images 4.

[0087] After a suitable neural network has been selected S114 in accordance with the previous disclosure, image data from the plurality of images 4 is input to the selected neural network. In one embodiment, only image data representing anatomical features that are depicted in all of the plurality of images 4 is input. In other words, image data of the plurality of images 4 representing anatomical features that are not depicted in all of the plurality of images 4 is not input to the neural network. One way to implement this selection of image data is to crop the images 31, 32, 34 to image regions 41, 42, 44 that include the anatomical features of all images and exclude all other anatomical features. The cropped images 41, 42, 44 are input to the selected neural network for processing.

[0088] By this method of analyzing the plurality of images 4 based on anatomical features and selecting a neural network that is trained on image data matching the anatomical features of the plurality of images 4, the chances of successfully re-identifying the person 38 in the plurality of images 4 as having the same identity are increased.

[0089] Proceeding to another embodiment, a further step of the method is to evaluate S109 the input vector before comparing S112 the input vector to the reference vectors. This is a quality assurance of the input vector with the purpose of maintaining a minimum level of success rate of the re-identification. The purpose is to filter out images of the plurality of images 4 that can lead to poor results from the neural network. The evaluation can comprise evaluating the input vector against preset quality conditions. The preset quality conditions can define that the input vector needs to represent at least one anatomical feature of a predefined list of anatomical features. The content of the predefined list can depend on the provided neural network, in particular, on which reference data the neural network has been trained on. For example, if the available neural network has been trained on reference data with different sets of anatomical features being shoulder, upper arm, elbow, forearm, and back of hand, the input vector can need to represent one of the anatomical features of elbow and hand for the plurality of images to qualify for use in the re-identification.

[0090] If the pre-set quality condition is met, the method continues at step S112 by comparing the input vector with the reference vector. If the pre-set quality condition is not met, the method can comprise a step S111 of discarding one or more images from the plurality of images 4.

[0091] A first example of a quality condition is that the input vector should have a minimum amount of anatomical features in it.

[0092] A second example of a quality condition is that the input vector should have a predetermined number of anatomical features from a predefined list. The predefined list can be related to the anatomical features the neural network has been trained on, to avoid processing a plurality of images with anatomical features the neural network has not been sufficiently trained on.

[0093] A third example of a quality condition is that the pose calculated from the anatomical features of the input vector should satisfy a certain condition. For example, the pose should correspond to a normal pose of the associated body part of the anatomical feature (in case of a human subject). The purpose of such a quality condition is to reduce the risk of performing the method on an image where the anatomical features in the input vector have been incorrectly estimated / determined for.

[0094] The discarding S111 of one or more images can comprise selecting which image or images to discard. The selection can be based on the anatomical features of the images. For example, if a first image lacks one or more anatomical features that are present in all other images of the plurality of images 4, the first image can be discarded. In the illustrated example, the first image can be the image 34 that lacks the anatomical features of the second eye that the rest of the images 31, 32 depict. The image 34 can thus be discarded, and the method can start again from the step S106 of determining anatomical features, now only based on the images 31 and 32 of the updated plurality of images 4.

[0095] It should be noted that the image sequences and plurality of images illustrated and discussed herein are provided as simplified examples, and are adapted to easily understand the inventive concept. In practice, the image sequences and plurality of images comprise more images. Typically, more than one object is detected in one or more images. The method can comprise selecting an object for an image in the plurality of images to perform the method on. Furthermore, the method can be adapted to compare the object of one image in the plurality of images with the objects of the other images in the plurality of images.

[0096] Figure 5An example of a plurality of images 5 captured by different cameras monitoring the same scene as previously described is illustrated, in which a person 38 is about to cross a road 39 on which a truck 37 is driving. In this example, the method can serve the purpose of evaluating whether the object 38 depicted in the images 51, 52 has the same identity. The images 51, 52 can be captured at the same point in time.

[0097] Figure 6 An example of a plurality of images 6 captured by different cameras monitoring different scenes is illustrated. The upper three images 61, 62, 63 forming a first image sequence correspond to a selection of images from Figure 3 The lower three images 64, 65, 66 forming a second image sequence depict two different objects 38, 68. Of course, the method does not know in advance whether the objects of the images have the same identity, e.g. if the object 68 of image 64 is the same person as the object 38 of image 63. Solving this problem is the actual purpose of the method.

[0098] According to the method, the objects 38, 68 are detected in the plurality of images 6. In the present embodiment, the plurality of images has been selected from the image sequences according to a temporal distance, i.e. there is a predetermined time gap between each of the images in each of the image sequences of the plurality of images 6. The method can comprise a further step of evaluating the selected plurality of images 6 and discarding images for which no object is detected. In this example, image 62 is discarded. The objects 38, 68 are detected in the remaining images 61, 63, 64, 65, 66 which now form the plurality of images 6. As mentioned above, the method can comprise a further step of selecting for an image the object to be compared with the objects of other images for the purpose of re-identification. The object 38 of image 61 can be selected to be compared with the object 68 of image 64, the object 38 of image 65 and the object 68 of image 66. The method can be performed on the group of images 61, 64, 65, 66 at the same time and can select to discard S111 one or more images if appropriate. Alternatively, the method can be performed on pairs of images of the group of images 61, 64, 65, 66. For example, first on the pair of images 61, 64, focusing on the object 38 of image 61 and the object 68 of image 64. This re-identification will likely result in a negative outcome, i.e. the object 38 in image 61 does not have the same identity as the object 68 of image 64. Next, image 61 can be compared with image 65, focusing on the object 38 of both images. This re-identification will likely result in a positive outcome, i.e. the object 38 in image 61 has the same identity as the object 38 of image 65. Alternatively, image 61 can be compared again with image 64, now focusing on the object 38 in image 64 (instead of object 68). This re-identification will likely have a positive outcome.

[0099] In other words, the method can be performed iteratively, where the plurality of images is updated in or before each iteration. Depending on the purpose of the re-identification, a different number of images is processed in one iteration. Irrespective of how many images and what the purpose of the re-identification is, the method relies on the inventive concept of selecting a neural network from a plurality of networks trained on different groups of anatomical features to perform the re-identification task based on a plurality of images depicting the object. As an example, it is to be understood that the present invention is not limited to the illustrated embodiments and that several modifications and variations are conceivable within the scope of the present invention.

[0100] To further help understanding the present invention, below is an overview of the claimed method and specific examples. The purpose of the present invention is to reduce the drawbacks of existing methods for object re-identification, i.e. the difficulty to re-identify an object based on images showing different numbers of anatomical features of the object. For example, some images depict a full-body object, while other images depict only an upper-body object. This drawback has been recognized by the inventors and exists for example in human objects. The inventors propose to establish several neural networks for object re-identification, where each network is trained on a different configuration of anatomical features of objects of the object class. Furthermore, the inventors propose to employ the neural network trained on the anatomical feature configuration that is most similar to the anatomical features depicted in all of the images in the group of images to be analyzed.

[0101] To not make the examples unnecessarily complex, we only provide two neural networks for object re-identification. Each neural network is trained on image data having a different group of anatomical features. Each group of anatomical features is represented by a keypoint vector, which is referred to as a reference vector. In this example, the keypoint vector is a one-dimensional binary vector, where each position in the vector indicates a certain anatomical feature. A vector position value of 1 means that the anatomical feature of that position is visible, while a value of 0 means that the anatomical feature is not visible. An example of such a keypoint vector can be as follows:

[0102] [a b c d e f]

[0103] The vector positions a-f indicate the following anatomical features:

[0104] a: eye

[0105] b: nose

[0106] c: mouth

[0107] d: shoulder

[0108] e: elbow

[0109] f: hand

[0110] For example, the keypoint vector [1 1 0 0 1] of a detected object in an image means that the eye, nose, mouth, and hand are visible, while the shoulder and elbow are not.

[0111] Each neural network is trained on image data having a different set of anatomical features. For example, a first neural network is trained on image data having a face, which includes a first set of anatomical features of an eye, a nose, and a mouth. A first reference vector representing the first set of anatomical features is [1 1 1 0 0 0]. A second neural network is trained on image data having a lower arm, which includes a second set of anatomical features of an elbow and a hand. A second reference vector representing the second set of anatomical features is [0 0 0 0 1 1].

[0112] These two neural networks can be described as neural networks trained to perform object re-identification based on different anatomical features in input image data. The first neural network is particularly good at performing object re-identification based on images depicting an eye, a nose, and a mouth, while the second neural network is particularly good at performing object re-identification based on images depicting an elbow and a hand.

[0113] Now to the input vector. This is also in the keypoint vector format. The input vector will be compared to the reference vectors in order to find the most similar reference vector and, thus, the most appropriately trained neural network for the task of object re-identification. To ease the comparison, the keypoint vector of the input vector can be constructed in the same way as the reference vectors, i.e. as [a b c d e f] above. However, comparing between keypoint vectors of different formats is a task easily solved by a person skilled in the art using conventional methods. For example, the input vector can have another size, i.e. more or fewer vector positions, and / or include more or fewer anatomical features. As long as it is clearly defined how to read out which anatomical features are detected and not from the keypoint vector, the comparison can be made.

[0114] However, we continue the discussion of the example of low complexity and construct the input vector in the same keypoint vector [a b c d e f] as the reference vectors. To determine the input vector, a plurality of images received are analyzed to determine which anatomical features are depicted in each of them. For anatomical features represented in all of the plurality of images, the respective vector position in the input vector is 1 and, thus, indicates that the anatomical feature is visible. For anatomical features not depicted in each of the plurality of images, the respective input vector position is 0, i.e. the anatomical feature is indicated as not visible. Let us assume we get the input vector [0 1 1 1 0 1], which means that the anatomical features of the nose, the mouth, the shoulder, and the hand are visible in each of the plurality of images.

[0115] Next, the input vector is compared to each of the reference vectors to determine the most similar reference vector according to a predefined condition. In other words, the input vector of [011101] is compared to each of [111000] and [000011]. The predefined condition can for example be the maximum number of overlapping anatomical features. The result of the comparison to the predefined condition is that the first reference vector [111000] is the most similar vector associated with the first neural network. Hence, the first neural network is selected to perform object re-identification based on the plurality of images with the aim to determine whether the plurality of objects depicted in the plurality of images have the same identity.

Claims

1. A method for object re-identification in an image of an object of an object type, the method comprising: providing a plurality of neural networks for object re-identification, wherein different ones of the plurality of neural networks have been trained on different sets of anatomical features for the object type to perform object re-identification, and wherein each set of anatomical features is represented by a reference vector in the form of a keypoint vector, wherein the keypoints represent the anatomical features, receiving a plurality of images of objects of the object type, determining keypoints representing anatomical features of the object type in each of the plurality of images, determining an input vector representing an anatomical feature determined in all of the plurality of images, wherein the input vector is in the form of a keypoint vector representing the anatomical feature, comparing the input vector with the reference vectors to determine the most similar reference vector according to predefined conditions, selecting a neural network from among the plurality of neural networks based on the comparison of the input vector and the reference vector, and Image data of a plurality of objects including all or part of the image data of the plurality of images are input to the neural network (#1) represented by the most similar reference vector to determine whether the plurality of objects have the same identity.

2. The method according to claim 1, wherein The object type is human.

3. The method according to claim 1, wherein The predefined condition defines that a reference vector equal to the input vector is determined as the most similar reference vector.

4. The method according to claim 1, wherein The predefined condition defines that a reference vector having the greatest overlap with the input vector is determined from among the reference vectors as the most similar reference vector.

5. The method according to claim 1, wherein The predefined condition defines determining, from the reference vectors, a reference vector comprising a maximum number of anatomical features defined by a priority list that overlap with the input vector.

6. The method according to claim 1, further comprising: evaluating the input vector with reference to a preset quality condition, When the preset quality condition is met, the step of comparing the input vector with the input image data is performed, and When the preset quality condition is not met, at least one image of the plurality of images is discarded, a new input vector is determined as the input vector based on the plurality of images, and the method is iterated from the step of evaluating the input vector.

7. The method according to claim 6, wherein: Said evaluating the input vector comprises comparing the input vector with a predefined list of anatomical features, at least one anatomical feature from said predefined list being supposed to be represented in the input vector.

8. The method according to claim 1, wherein The multiple images are captured by one camera at multiple points in time, by multiple cameras covering the same scene from different angles, or by multiple cameras depicting different scenes.

9. The method according to claim 1, wherein The inputting of the image data of the plurality of images includes inputting image data representing only the anatomical features depicted in all of the plurality of images.

10. The method according to claim 1, wherein The step of receiving the plurality of images comprises: Images are captured by one or more cameras, and Different images are selected to form the plurality of images based on a predetermined frame distance, time gap, image sharpness, pose of a depicted object, resolution, aspect ratio of a region, and plane rotation. 11 . A non-transitory computer-readable recording medium having computer-readable program code recorded thereon, the computer-readable program code being configured to perform the method of claim 1 when executed on a device having processing capabilities.

12. A controller for controlling a video processing unit to facilitate object re-identification, the controller having access to a plurality of neural networks for object re-identification in images of objects of an object type, wherein Different ones of the plurality of neural networks have been trained on different sets of anatomical features for the object type to perform object re-identification, and wherein each set of anatomical features is represented by a reference vector in the form of a vector of keypoints, wherein the keypoints represent the anatomical features, the controller comprising: a receiver configured to receive a plurality of images of an object of the object type; a determining means configured to determine keypoints representing anatomical features of the object type in each of the plurality of images (4) and to determine an input vector representing the anatomical features determined in all of the plurality of images, wherein the input vector is in the form of a keypoint vector representing the anatomical features, a comparison component adapted to compare the input vector with the reference vector to determine the most similar reference vector according to a predefined condition, a selection function that selects a neural network from among said plurality of neural networks based on said comparison of said input vector and said reference vector, an input component configured to input all or part of the image data of a plurality of objects including the image data of the plurality of images into the neural network represented by the most similar reference vector to determine whether the plurality of objects have the same identity, and A control component is configured to control whether the video processing unit regards the multiple objects as having the same identity.

Citation Information

Patent Citations

  • Pedestrian re-identification method and a related product

    CN109657533A