Method and device for anonymizing re-identification data in visual tracking system
By introducing anonymization layer and tracking permission tokens in the visual tracking system, anonymizing and distributing re-identification data, the problem of data propagation and tracking range limitation in the visual tracking system is solved, and effective control of data anonymization and tracking range is achieved.
Patent Information
- Application Number
- CN202411839393.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-12-13
AI Technical Summary
In existing visual tracking systems, the dissemination of re-identification data may lead to unnecessary dissemination of personal data, and it is difficult to effectively limit the scope of visual tracking, resulting in unauthorized tracking.
By introducing anonymization layer in the visual tracking system, the reidentification data is anonymoused using predefined one-way functions, and the anonymous reidentification data items are distributed to authorized tracking clients through the tracking permission token, ensuring data anonymization and limiting of the tracking scope.
The anonymization of visual tracking data is achieved, preventing unauthorized data access, ensuring effective restrictions on the tracking scope, and protecting the security of personal data.
Smart Images

Figure CN120182332A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of visual tracking systems. In particular, the present disclosure presents a novel method and apparatus for anonymizing re-identification data in a visual tracking system, as well as a novel method and apparatus for performing visual tracking using anonymized re-identification data. Background Art
[0002] Recent technological developments in the field of object re-identification have provided more powerful methods for visually tracking targets such as people or objects within the field of view (FOV) of one camera or across the FOVs of several cameras. This is very useful in applications such as object search and loitering detection. According to an established and widely practiced paradigm, visual tracking is not performed directly on image data, but on so-called re-identification data (reID data). As reID data, feature vectors that each represent the visual appearance of a tracking target are typically used. The feature vectors can be described as a low-dimensional representation of the image data, which on the one hand saves the amount of data to be processed and on the other hand allows for precise matching.
[0003] Visual tracking can be performed by a tracking client that is separate from the camera that acquires the original image data. In this case, the image data must be transmitted from the camera to the tracking client via a data connection, which raises concerns about the unnecessary dissemination of personal data (e.g., via eavesdropping). If visual tracking is to be performed across the FOVs of several cameras, the transmission of image data is usually inevitable. If the camera supplies reID data instead of image data to the tracking client, the concerns about dissemination are equally valid, since the reID data can enable unauthorized tracking of previously identified targets.
[0004] Given these concerns, it is desirable to preprocess or modify the re-identification data at the source such that the scope of visual tracking is bounded with respect to space and / or time. As an example of spatial delimitation, a system owner may want to specify that re-identification data from one image source is only to be used for visual tracking within the FOV of that image source.
[0005] In another example, the system owner may agree to use this re-identification data in combination with re-identification data from a second image source, but not to use it in combination with re-identification data from a third image source; in other words, the allowed use of the re-identification data spans the FOVs of the first and second image sources. In the case where an unauthorized party (eavesdropper) gains access to the re-identification data, the technical mechanism should cause their tracking attempt to fail. Summary of the Invention
[0006] One object of the present disclosure is to provide methods and apparatuses for defining the scope of visual tracking relative to space and / or time. (Here, if the scope is restricted to a plurality of specified image sources, even if the spatial positions of the image sources are unknown, the scope can be said to be defined relative to space.) A further object is to provide methods and apparatuses for anonymizing the data on which visual tracking is based. A further object is to provide methods and apparatuses for providing re-identification data that supports re-identification only within the FOV of a specific group of image sources. A further object is to provide methods and apparatuses for protecting anonymity. Yet a further object is to propose methods and apparatuses for performing visual tracking using anonymized data.
[0007] At least some of these objects are achieved by the present invention as defined in the independent claims. The dependent claims relate to advantageous embodiments of the present invention.
[0008] In a first aspect of the present disclosure, a visual tracking system is provided having the following main components:
[0009] A plurality of image sources,
[0010] An anonymization layer, and
[0011] A plurality of tracking clients.
[0012] The image sources are configured to provide images within their respective FOVs, wherein each image source is configured to detect, in the image obtained from the image source, a respective sub-region m that contains a tracking target i , and to compute, for each sub-region, a feature vector f(m i ) that represents the visual appearance of the tracking target therein. Further, the anonymization layer is implemented in a processing circuit separate from the tracking clients, such as in a processing circuit located at the same location as a respective one of the image sources and / or in a processing circuit belonging to a coordination entity in the visual tracking system. The anonymization layer is configured to provide a first re-identification data item g i,1 = h([f(m i ), σ1]) by anonymizing each feature vector using a predefined one-way function h modified by a first tracking permission token σ1. The first tracking permission token σ1 is specific to a first group of image sources. The anonymization layer is further configured to provide a second re-identification data item g i,2 = h([f(m i ), σ2]) by anonymizing each feature vector using the predefined one-way function h now modified by a second tracking permission token σ2. The second tracking permission token σ2 is specific to a second group of image sources and is different from the first tracking permission token, σ2 ≠ σ1. The anonymization layer is further configured to disclose to the tracking clients the position X(m of the corresponding sub-region m i i )'s first-level recognition data items and second-level recognition data items, while preventing access to the feature vectors. The first set of image sources includes at least one image source in the image sources. The second set of image sources includes at least one image source in the image sources. Finally, the tracking client is configured to perform re-identification on the tracking target using the data obtained from the anonymization layer. Among the tracking clients, there are at least: a first tracking client authorized to perform re-identification in the FOV of the first set of image sources and obtain the first-level recognition data items with location annotations, and a second tracking client authorized to perform re-identification in the FOV of the second set of image sources and obtain the second-level recognition data items with location annotations. The tracking clients can obtain them by receiving the re-identification data items on an internal or external data connection, or by retrieving the re-identification data items from a shared memory.
[0013] Advantageously, since the anonymization layer prevents access to the feature vectors (and the anonymization layer is separated from the tracking client), and since the first-level recognition data items and the second-level recognition data items are generated by different tracking permission tokens σ1, σ2 specific to the first set of corresponding image sources / the second set of image sources, each tracking client is restricted to performing re-identification in the FOV of its corresponding image source group. If it is assumed that the first tracking client attempts to perform re-identification in a set that includes both the first-level recognition data items and the second-level recognition data items, then the first tracking client will never be able to find a match between one of the first-level recognition data items and one of the second-level recognition data items. Thanks to the collision-free property of the one-way function, even if the function h is applied to the same feature vector f(m0), the uniqueness of the tracking permission tokens σ1≠σ2 will ensure
[0014] g 0,1 = h([f(m0), σ1]) ≠ h([f(m0), σ2]) = g 0,2 (1)
[0015] Therefore, it is technically meaningless for the first tracking client to attempt and go beyond the first set of image sources. For general re-identification data (e.g., assuming the original feature vector has been used as re-identification data), there is no inherent mechanism to prevent the first tracking client from performing re-identification outside the FOV of the first set of image sources.
[0016] Again due to the one-way property of the function h, the feature vector f(m i ) is anonymized. The downstream parties of the anonymization layer cannot reconstruct the feature vector f(m i,1 , g i,2 from one of the re-identification data items g i ). Even with full access to a large set of re-identification data, inverting the function h is a computationally infeasible task.
[0017] In a second aspect of the present disclosure, a method is provided for providing anonymized data to facilitate re-identification in a visual tracking system. The method includes: detecting in images obtained from a plurality of image sources sub-regions m each containing a tracking target i ; for each sub-region, computing a feature vector f(m i ) representing the visual appearance of the tracking target therein. For a first subgroup of the image sources, providing a first re-identification data item g i,1 = h([f(m i ), σ1]) by anonymizing each feature vector using a predefined one-way function h modified by a first tracking permission token σ1; and disclosing to a first tracking client the first re-identification data item labeled with the location X(m i ) of the corresponding sub-region. For a second subgroup of the image sources, providing a second re-identification data item g i,2 = h([f(m i ), σ2]) by anonymizing each feature vector using a predefined one-way function h modified by a second tracking permission token σ2 different from the first tracking permission token; and disclosing to a second tracking client the second re-identification data item labeled with the location X(m i ) of the corresponding sub-region. The method further includes preventing access to the feature vectors.
[0018] The method according to the second aspect facilitates (i.e., assists, supports, enables) re-identification in a visual tracking system because it provides re-identification data items that tracking clients will search for matches among to track the tracking target. As explained above, the proposed method for generating re-identification data items also ensures anonymity. Further, thanks to the uniqueness of the tracking permission tokens, the re-identification data items are generated in a way that bounds the scope of visual tracking to a specified subgroup of the image sources. Multiple instances of the method according to the second aspect can be executed such that the first tracking permission token σ1 is specific not only to the first subgroup of the image sources but also to a first group of image sources including the first subgroup, and / or such that the second tracking permission token σ2 is specific not only to the second subgroup of the image sources but also to a second group of image sources including the second subgroup. For example, the first group of image sources may consist of the first subgroup and a further subgroup, and an independent process uses the same first tracking permission token σ1 to generate re-identification data items for the further subgroup.
[0019] Some steps of the method according to the second aspect can be executed in an anonymization layer, and some steps can also be executed in the image sources. In the implementation of the method, the steps of detecting the sub-regions m i and computing the feature vectors f(m i ) for each sub-region can be delegated to the image sources. Accordingly, the anonymization layer does not need to execute more than the following steps:
[0020] - For a first subgroup of image sources, provide first re-identification data items by anonymizing each feature vector using a predefined one-way function h modified by a first tracking permission token; and disclose the first re-identification data items labeled with the locations of the corresponding sub-regions to a first tracking client.
[0021] - For a second subgroup of image sources, provide second re-identification data items by anonymizing each feature vector using a function h modified by a second tracking permission token; and disclose the second re-identification data items labeled with the locations of the corresponding sub-regions to a second tracking client.
[0022] - Prevent access to the feature vectors.
[0023] It should be understood that the anonymization layer may have initially obtained feature vectors calculated for sub-regions in images from multiple image sources, where each sub-region contains a tracking target, and the feature vectors represent the visual appearance of the tracking target therein.
[0024] According to a third aspect of the present disclosure, a method for visually tracking a tracking target in the field of view of a first set of image sources is provided. The method includes: receiving first re-identification data items g i,1 ; searching for matching re-identification data items among the first re-identification data items; and for a set of mutually matching re-identification data items, tracking a tracking target based on the locations labeled by the mutually matching re-identification data items. According to the third aspect, each first re-identification data item is derived from a feature vector f(m i ), the feature vector f(m i ) represents the visual appearance of the tracking target in a sub-region of an image obtained from an image source in the first set, and each first re-identification data item is labeled with the location of the corresponding sub-region. Further, all first re-identification data items are calculated using a predefined one-way function h modified by a first tracking permission token σ1 specific to the first set of image sources.
[0025] Although the re-identification data is provided in the form of anonymized data, the method according to the third aspect still enables visual tracking. In particular, the method may include searching for matching re-identification data items among a set of re-identification data items that have been calculated using a single predefined one-way function h modified by a single tracking permission token σ1. (As explained above, the input to this calculation is feature vectors that each represent the visual appearance of a detected tracking target.) That is, the method excludes searching for matching re-identification data items among re-identification data items calculated using different one-way functions and / or among re-identification data items calculated using one-way functions modified by two or more different tracking permission tokens.
[0026] In a fourth and fifth aspect of the present disclosure, there is provided an apparatus or a cluster of apparatuses including processing circuitry configured to perform the method according to the second aspect (providing anonymized data to facilitate re-identification) or the method according to the third aspect (performing visual tracking).
[0027] The present disclosure further provides a computer program comprising instructions for causing a computer to perform the above-described methods. The computer program may be stored or distributed on a data carrier. As used herein, a "data carrier" may be a transient data carrier (such as a modulated electromagnetic wave or light wave) or a non-transient data carrier. Non-transient data carriers include volatile and non-volatile memories such as permanent and non-permanent storage media of magnetic, optical, or solid-state types. Still within the scope of "data carrier", such memories may be fixedly installed or portable.
[0028] Some embodiments also define the scope of visual tracking with respect to time, i.e., after the expiration of the validity period of the first tracking permission token, the first tracking permission token σ1 is stopped being used (in the anonymization layer). Thereafter, the anonymization layer may anonymize the feature vectors using a predefined one-way function h modified by an alternative first tracking permission token σ′1. Thus, because the one-way function provides non-conflicting output values, the first tracking client will never find a match between the re-identification data items generated before and after the token replacement, even for the same feature vector (i.e., similar to (1), h([f(m0),σ1])≠h([f(m0),σ′1])).
[0029] To define the scope of visual tracking only with respect to time, the same tracking permission token is used to modify the one-way function h for all image sources in the visual tracking system. When the validity period of the tracking permission token expires, the one-way function h is alternatively modified by the tracking permission token (still for all image sources in the visual tracking system). This allows the tracking client to perform re-identification in the FOV of all image sources in the visual tracking system, but one validity period at a time.
[0030] For the purposes of the present disclosure, the term "re-identification data" (or reID data) refers to the quantity or variable that forms the basis of the re-identification process, which is a matching process that identifies image data from different times and locations as pointing to the same tracking target (such as a person or an object). The reID data can be a proxy for the actual image data (such as a feature vector derived from the image data). In an individualized form, the feature vector can be referred to as an reID data item. The reID data can be provided in the form of non-anonymized data (e.g., a feature vector) or anonymized data (e.g., a hash value of a feature vector). According to the first, second, and further aspects of the present disclosure, the reID data should be provided as anonymized data. Hashing and other types of anonymization processes may result in a change in the data type; for example, the hash value of a vector can be a scalar, although there are special hashing techniques that return a vector.
[0031] As used herein, the "location" X(m i ) of sub-region m i can refer to the location of the sub-region in the image or the geographical location. The geographical location can correspond to the location of the image source from which the image was acquired, which is independent of the location of the sub-region in the image. Further, the geographical location can be an approximate location of the portion of the FOV of the image source depicted by sub-region m i .
[0032] In general, all terms used in the claims should be interpreted according to their ordinary meaning in the technical field, unless otherwise explicitly defined herein. All references to "a / an / the element, apparatus, component, manner, step, etc." should be construed broadly as referring to at least one instance of the element, apparatus, component, manner, step, etc., unless otherwise explicitly stated. The steps of any method described herein need not be performed in the exact order disclosed, unless explicitly stated. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Aspects and embodiments are now described by way of example with reference to the accompanying drawings, in which:
[0034] Figure 1A and Figure 1B show two visual tracking systems;
[0035] Figure 2 is a flowchart of a method for providing anonymized data to facilitate re-identification in a visual tracking system;
[0036] Figure 3 is a flowchart of a method for visually tracking a tracking target based on anonymized re-identification data; and
[0037] Figure 4Illustrate the generation of anonymized re-identification data from the acquired images. Detailed implementation
[0038] Aspects of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, which show certain embodiments of the invention. However, these aspects may be embodied in many different forms and should not be construed as limiting; rather, these embodiments are provided by way of example so that the present disclosure will be thorough and complete and will fully convey the scope of all aspects of the invention to those skilled in the art. Throughout the specification, the same numbers refer to the same elements.
[0039] Figure 1A A visual tracking system 100 suitable for visually tracking tracking targets P1, P2, P3 is shown. The tracking targets P1, P2, P3, which can be a person or an inanimate object, are exemplified as vehicle P1 and two different people P2, P3 in Figure 1A Conceptually, visual tracking technology can be said to be based on the assumption that based on the unique or substantially unique visual features within the visual tracking range, the tracking targets can be identified with high accuracy. In the case of a person, the unique visual features may include height, posture, body shape, facial features, and the combination of clothing and shoes. In the case of a vehicle, the license plate may be preferred. For each tracking target, the possible output of visual tracking can be a list of positions and an indication of when the tracking target appears at these positions.
[0040] Figure 1A The visual tracking system 100 in includes four image sources 120.1, 120.2, 120.3, 120.4 arranged to provide corresponding FOVs 129.1, 129.2, 129.3, 129.4 and images of at least two tracking clients 140.1, 140.2. Each tracking client 140 includes a processing circuit 141 and a memory 142, wherein the other memory 142 is adapted to store a computer program 143 executable by the processing circuit 141, as well as the input data and output data of the re-identification process, and other objects. The visual tracking system 100 further includes an anonymization layer 130; the anonymization layer 130 is a virtual entity implemented in a coordination entity 150. The coordination entity 150 can be a centralized hardware component or a centralized software process in the visual tracking system 100. The coordination entity 150 can be implemented in a dedicated server or host in a computer network, or implemented as a process running on a computer or virtual machine having additional responsibilities in the visual tracking system 100. The coordination entity 150 runs on a processing circuit 151 and has a memory 152 at its disposal; the functions of the coordination entity 150 can be encoded as a computer program 153.
[0041] The image source 120, the coordination entity 150, and the tracking client 140 are linked by a wired or wireless data connection (e.g., by connecting to a common data network). As will be explained below, the data connections from the image source 120 to the coordination entity 150 should preferably be protected from unauthorized parties because they convey feature vectors that have not yet undergone anonymization. (These data connections are also sensitive in such embodiments of the visual tracking system 100 where the feature vectors are computed in the coordination entity 150, i.e., because the data connections convey image data.) The data connections can be protected from eavesdropping and other attacks by unauthorized parties through appropriate end-to-end encryption, tunneling, or by allowing the data connections only through physically protected wired lines that are reasonably invulnerable to intrusion, such as the internal bus of a device.
[0042] In Figure 1A the visual tracking system 100, the novel techniques proposed herein are used to spatially delimit the scope of visual tracking, and more specifically, to enable the first tracking client 140.1 to perform visual tracking in the FOV of the first set 110.1 that includes the first image source 120.1 and the second image source 120.2, and to enable the second tracking client 140.2 to perform visual tracking in the FOV of the second set 110.2 that includes the third image source 120.3 and the fourth image source 120.4. For this purpose, the first tracking client 140.1 is supplied with the first re-identification data item g 1,1 , g 2,1 , g 3,1 , …, and the second tracking client 140.2 is supplied with the second re-identification data item g 1,2 , g 2,2 , g 3,2 , …. The first set 110.1 of image sources and the second set 112.2 of image sources may be non-overlapping, or they may at least partially overlap.
[0043] Without departing from the scope of the present disclosure, re-identification data originating from the same set of image sources can be provided to more than one tracking client. As Figure 1A prompted by the dashed box in the upper right corner, the third tracking client 140.3 and the fourth tracking client 140.4 can perform re-identification in the FOV of the second set 110.2 of image sources based on the second re-identification data item with location annotation.
[0044] The functions of the different components of the visual tracking system 100 will now be described in more detail. The image sources 120.1, 120.2, 120.3, 120.4 include lenses, photosensitive components (image sensors), and image processing components, through which they provide digital image data representing still images or video sequences of their respective FOVs 129.1, 129.2, 129.3, 129.4. The image source 120 can be, for example, a digital video camera.
[0045] In Figure 1A the example shown, each image source 120 further includes processing circuitry 121 configured to detect sub-regions m1, m2, …, m6 in the image M such that each sub-region contains a tracking target. Figure 4 The upper part of i shows an example where this detection process is applied to four images M; visually observable, sub-regions m1, m4, m6 contain the same person at different positions (and thus, at different time points), sub-region m5 contains a different person, and sub-regions m2, m3 contain the same vehicle. The detection can be based on object detection algorithms or object classification algorithms, such as body part detection algorithms, face detection algorithms, or in the case of vehicles, based on algorithms configured to detect alphanumeric characters or visual features (e.g., on license plates) of the vehicle. In different implementations, the algorithm can be configured to output bounding boxes, or separate bounding box post-processing can be performed in other ways downstream of the algorithm, thereby obtaining sub-regions m1, m2, …, m6. In terms of notation, each sub-region m
[0046] Figure 1A the processing circuitry 121 of the image source 120 in
[0047] aAPjFcw7bOw74uD4GEVdpo5v0-Sal1eoguVkYuwc1srRbb0OFhLOVOwUPA,
[0048] It represents 128 bytes, with 3 bits per byte. In the prior art, the feature vector, together with an appropriate distance metric, is the basis for matching in the re-identification process, i.e., evaluating whether two feature vectors are close enough with respect to the distance metric such that they should be considered relevant to the same tracking target. The following references review various person re-identification methods within this framework:
[0049] · M. Ye et al., "Deep Learning for Person Re-Identification: A Survey and Outlook," arXiv preprint, arXiv:2001.04193 (2021);
[0050] · Zahra et al., "Person Re-Identification: A Review of Domain-Specific Open Challenges and Future Trends," Pattern Recognition, Vol. 142 (2023), 109669, DOI: 10.1016 / j.patcog.2023.109669; and
[0051] · L. Zheng et al., "Person Re-Identification: Past, Present, and Future," arXiv preprint, arXiv:161002984 (2016).
[0052] How these results can be used as the basis for re-identification in the specific visual tracking system 100 described herein will be explained below.
[0053] Broadly speaking, a feature vector can be described as a low-dimensional representation of the visual appearance of a tracking target. The feature vector can be, for example, a license plate number representing alphanumeric characters read from an image of a vehicle's license plate (e.g., f(m1) = 'ABC123'). In this example, according to prior art re-identification techniques, a match may require all characters to be equal, otherwise the vehicle will not be recognized as the same.
[0054] In another example, the feature vectors are calculated based on transparent (“handcrafted”, human-defined) definitions (such as color- or texture-based features (e.g., weighted color histograms, maximally stable color regions, recurrent highly structured patches)) or attribute-based features (e.g., clothing, biometrics). This definition of the feature vectors is independent of whether the object has an artificial label consisting of a license plate. The calculation of the feature vectors can further take into account aspects such as the appearance from multiple viewpoints or the motion pattern of the object that can be derived from the video data. The feature vectors can be low-dimensional in the sense that they include a small number of components and / or each component takes values in a set with a finite cardinality (e.g., by rounding real numbers to integers). In this case, according to re-identification techniques of the prior art, the matching of two feature vectors x = (x1, x2, …, x N ), y = (y1, y2, …, y N ) can correspond to the exact equality of the vector components:
[0055] x n = y n , 1 ≤ n ≤ N, (2)
[0056] Or it can correspond to a separation with respect to a distance metric up to a threshold D0
[0057] d(x, y) < D0 (3)
[0058] The distance metric in (3) can correspond, for example, to an l p distance:
[0059]
[0060] where p > 0 is a predefined constant. Additionally, the feature vectors can be compared with a distance metric d(x, y) calculated by a machine learning model that is trained to mimic correct or human-like re-identification decisions; then, a closed-form definition of the distance metric may not exist and is not required either.
[0061] In a third example, the feature vectors are computed by a machine learning model, such as a convolutional neural network (CNN), that is trained to mimic correct or human-like re-identification decisions. The training can include supervised or unsupervised learning. When the feature vectors are computed in this way, the computation does not follow any transparent definition (which is not necessary for the teachings herein anyway), and their results generally cannot be explicitly interpreted (e.g., in terms of the appearance aspects of the tracking target). Further, the computation can only be faithfully repeated by the machine learning model itself. Depending on how the machine learning model is trained, the feature vectors computed in this way can be compared against closed-form distance metrics (like Equation (2) and Equation (3)), or against distance metrics computed by the machine learning model.
[0062] In Figure 4 , reference numeral 401 denotes the process of providing feature vectors f(m1), f(m2), …, f(m6) based on image M. In Figure 1A the exemplary visual tracking system 100 illustrated in
[0063] A further process 402 determines the positions X(m1), X(m2), …, X(m6) of the sub-regions m1, m2, …, m6. Each position can refer to the position of the sub-region m i in the image M (e.g., represented in pixel coordinates or as a percentage of the FOV size), or it can be a geographical location (e.g., represented in a global reference system such as WGS84). For the sake of clarity, if the position is represented in pixel coordinates, local geographical coordinates, or other local coordinates, the position should preferably include a direct or indirect identifier of the image source. The geographical location can correspond to the (fixed) position of the image source 120 from which the image was acquired, which is independent of the position of the sub-region in the image. Further, the geographical location can be an approximate position of a portion of the FOV of the image source depicted by the sub-region m i . Thus, at least for some definitions of the positions X(m i ), the process 402 can be implemented in the image source 120 or in another component of the visual tracking system 100 (such as in the coordination entity 150), provided that the other component has access to the relevant information (such as the geographical location of the image source 120 or the geographical locations of different portions of the FOV of the image source 120 that may have been pre-estimated). These positions X(m1), X(m2), …, X(m6) can be used as components of the anonymization layer 130.
[0064] Anonymization layer
[0065] The functionality of the anonymization layer 130 will now be described in accordance with method 200 illustrated by the flowchart in Figure 2 , where the initial steps 210 and 211 correspond to processes 401 and 402 in Figure 4 . As mentioned, they may be performed in the image source 120 (or co-located therewith), in the coordination entity 150, or even in different components of the visual tracking system 100. The initial steps 210 and 211 may be performed in or upstream of the anonymization layer 130. Thus, in some embodiments, the anonymization layer 130 may be configured to perform only steps 212, 213, and 214. Figure 1A The architecture shown in i where the anonymization layer is centralized in the coordination entity 150 may avoid the need to distribute tracking entitlement tokens to multiple parallel processes and / or may simplify the replacement of tracking entitlement tokens when their validity period expires. It should be understood that when steps 210 and 211 are performed in the image source 120 (or co-located with the image source 120), they may correspond to multiple processes that are executed in parallel; for example, there may be a process of detecting sub-region m i and calculating the feature vector f(m i ) for each image source.
[0066] Steps 212 and 213 are performed once for each subgroup 111 of the image source and may thus be repeated for two or more subgroups 111. For each execution of steps 212 and 213, a different tracking entitlement token should be used. In an embodiment where the anonymization layer 130 is not implemented in the processing circuitry 151 of the coordination entity 150 but is distributed across multiple processors 121.1, 121.3, 121.4 that are co-located with a respective one of the image sources 120.1, 120.3, 120.4 (see Figure 1B ), it may occur that two instances of method 200 (executed on two different processors) use the same tracking entitlement token σ1 with respect to two different subgroups 111.1, 111.2 of the image source. In this setting, the tracking entitlement token σ1 may be considered specific to the image source group 110.1 that is the union of the two subgroups 111.1, 111.2.
[0067] For the first subgroup 111.1 of the image source, step 212 includes providing 212.1 a first re-identification data item by anonymizing each feature vector using a predefined one-way function h, where the one-way function h is modified by the first tracking entitlement token σ1 as follows:
[0068] g i,1 = h([f(m i ), σ1]), i ∈ I1, (4)
[0069] Among them, I1 is the first index set. The tracking permission token σ1 can be represented as a bit string. The symbol [·,·] refers to a combination operation such as string concatenation. The quantity (4) can have the appearance of a single number, a bit string, an alphanumeric string, or a vector, etc.
[0070] Assume that the one-way function is irreversible and there are no obvious collisions. As a one-way function, a hash function can be used, especially an encryption hash function that is considered to provide a sufficient level of security given the sensitivity of the image data. Examples are SHA-256, SHA3-512, RSA-1024, and possibly MD5. When applied as shown in (4), the tracking permission token can be regarded as acting as an encryption salt for the one-way function; conceptually, it defines the modified one-way function H(x) = h([x,σ1]). Because of the irreversibility of the one-way function h - even knowing the tracking permission token σ1 - it is not always necessary to treat the tracking permission token σ1 as a secret. Further, the re-identification data item g 1,1 ,g 2,1 ,g 3,1 ,… receivers (such as the tracking client 140) can use the re-identification data item without knowing the value of the tracking permission token σ1.
[0071] As explained above, the feature vector is an example of re-identification data that can be used in the re-identification process. Several useful definitions of the feature vector are known in the art, and according to condition (3), they can be classified based on whether equation (2) is required to determine a match or whether a non-zero separation less than the threshold D0 should also be accepted as a match. The re-identification data item g i,1 is also a kind of re-identification data that can be used in the re-identification process that basically conforms to the latest technology, that is, no drastic adjustments are required on the receiving side.
[0072] If the underlying feature vector is defined as requiring equation (2), the receiver can directly use the re-identification data item g 1,1 ,g 2,1 ,g 3,1 ,… This is because due to the conflict-free nature of h, the equation g j,1 = g j′,1 will imply that the underlying feature vectors are also equal, f(m j ) = f(m j′ ).
[0073] If instead the underlying feature vectors are defined such that non-zero separations less than a threshold D0 are accepted as matches (condition (3)), then a proximity-preserving one-way function h (a proximity-preserving hash function) should be used. A function h is proximity-preserving if for all feature vectors x, y such that d(x, y) < D0, it maintains d(h(x), h(y)) < D′0, where D′0 is a constant. Stated without formulas, for any pair of feature vectors closer than the threshold D0, there is a uniform bound D′0 on the distance of the anonymized feature vectors (i.e., the re-identified data items). Specifically, the proximity-preserving one-way function h should be such that for some constant D′0, d(h([x, σ]), h([y, σ])) < D′0, where σ is the tracing entitlement identifier. (The distance function may have to be defined differently when applied to re-identified data items since they have different formats or data types, but this is implicit here so as not to unduly burden the notation.) Using a proximity-preserving one-way function will allow tracing clients further downstream in the processing chain to search for matches based on a modified version of the proximity test (3), namely:
[0074] d(g j,1 ,g j′,1 ) < D′0. (3′)
[0075] The proximity-preserving property can be achieved by partitioning the feature vectors into shorter sub-vectors (each having one or more components) and hashing each sub-vector separately; in such an implementation, the re-identified data items include multiple sub-hashes of the sub-vectors (salted by the tracing entitlement tokens). If the feature vectors are partitioned into sub-vectors each having one component and hashed separately, the resulting re-identified data items can be a vector of the same length as the feature vectors. Further proximity-preserving hash functions are described in "Spectral Hashing" by Y. Weiss et al. in "Advances in Neural Information Processing Systems 21" (NIPS 2008), edited by D. Koller et al., ISBN 9781605609492:
[0076] re-identified data item g 1,1 ,g 2,1 ,g 3,1, … are labeled with corresponding positions X(m1), X(m2), X(m3), … and are disclosed to the tracking clients permitted to perform visual tracking within the FOV of the first subgroup 111.1 of the image source in step 213.1. Disclosing the re-identification data items to the tracking clients may include sending them to the tracking clients in a message or storing them in a shared memory to which the tracking clients have access. The shared memory may be configured as a publish / subscribe (Pub / Sub) messaging service. In this example, the first tracking client 140.1 is permitted to perform visual tracking within the FOV of the first subgroup 111.1 of the image source. Labeling the re-identification data items does not necessarily imply any modification to the re-identification data items themselves; rather, labeling may be achieved by storing the re-identification data items and their positions in a common data structure (such as Figure 4 as shown), in associated fields of a table, and / or by creating a pointer from the position to the re-identification data item or another computer-readable association, and vice versa.
[0077] For the second subgroup 111.2 of the image source, step 212 includes providing 212.2 first re-identification data items by anonymizing each feature data item using a predefined one-way function h modified by the second tracking permission token σ2, as follows:
[0078] g i,2 = h([f(m i ), σ2]), i ∈ I2, (5)
[0079] where the index sets I1, I2 may be disjoint or have non-zero overlap. These re-identification data items g 1,2 , g 2,2 , g 3,2 , … are labeled with corresponding positions X(m1), X(m2), X(m3), … and are disclosed to the tracking clients permitted to perform visual tracking within the FOV of the second subgroup 111.2 of the image source. In this example, the second tracking client 140.2 and optionally the third tracking client 140.3 and the fourth tracking client 140.4 are permitted to perform visual tracking within the FOV of the second subgroup 111.2 of the image source.
[0080] In Figure 4 , the provision of the first and second re-identification data items corresponds to process 403.1 and process 403.2, respectively. Annotating the second re-identification data items with positions is illustrated as process 404. There is a corresponding process (not shown) for annotating the first re-identification data items with positions.
[0081] Method 200 further includes step 214 of preventing access to the feature vectors. Measures will be taken to prevent any party from inspecting or accessing the feature vectors, or to make such attempts very difficult. To prevent access to the feature vectors, as explained above, the feature vectors may be transmitted over a data connection that is protected from eavesdropping and other attacks by unauthorized parties. Further, when storing or caching the feature vectors, a sufficiently protected memory may be used. Note that, at least in some embodiments, step 214 is limited to preventing access to those feature vectors controlled by the anonymization layer 130, e.g., feature vectors stored in the memory of a component acting as the anonymization layer. In such an embodiment, other components of the visual tracking system 100 may be responsible for preventing access outside the scope of control of the anonymization layer 130, i.e., protecting access to the data connection over which the feature vectors are transmitted upstream of the anonymization layer 130.
[0082] Tracking client
[0083] Go to Figure 3 , and now the behavior of the individual tracking client 140 will be described according to method 300. It is recognized that the same operations may be performed in a general-purpose processor. Method 300 will be described from the perspective of a tracking client that is authorized to visually track a tracking target in the FOV of a first set 110.1 of image sources (i.e., for which a first tracking permission token σ1 has been used to generate re-identification data items). In the running example, this corresponds to the first tracking client 140.1.
[0084] In an initial step 310, the tracking client receives a first re-identification data item g i,1 . This may include receiving the first re-identification data item g i,1 in a message or retrieving them from a shared memory. In fact, since the scope of visual tracking is spatially bounded thanks to the non-matchability (1), it is acceptable to make the first re-identification data item g i,1 and any second re-identification data item g i,2 , third first re-identification data item, etc. available in the same shared memory. Each of the first re-identification data items is derived from a feature vector f(m i ), the feature vector f(m i ) representing the visual appearance of the tracking target in a sub-region m i of an image obtained from an image source in the first set, and each first re-identification data item is labeled with the position X(m i)。As explained above, using the predefined one-way function h modified by the first tracking permission token σ1 of the first group specific to the image source more precisely provides the first re-identification data item. To execute the method 300, it is not necessary for the tracking client to confirm that the first re-identification data item has been provided in this particular way; the fact that the re-identification data item is an anonymized feature data item may even be opaque to the tracking client.
[0085] In step 311, the tracking client then continues to search for a matching re-identification data item among the received first re-identification data items g i,1 .
[0086] If the underlying feature vector f(m i ) is defined as discrete values or otherwise (e.g., projection onto a low-dimensional subspace, rounded to integer values) such that it is considered a match only when equation (2) is satisfied, then the re-identification data item also matches only when all of its components are equal. That is, the tracking client searches the set G P1 of first re-identification data items such that for all data item pairs g j,1 , g j′,1 ∈ G P1 the equality condition g j,1 = g j′,1 is satisfied. Since the one-way function h modified by the first tracking permission token σ1 is collision-free, the equality condition implies that the underlying feature vectors are also equal, f(m j ) = f(m j′ ). Each such set G P1 , G P2 , G P3 can be considered to correspond to a tracking target P1, P2, P3.
[0087] If instead, the underlying feature vector is defined according to the above criterion (3) such that a non-zero separation less than the threshold D0 is accepted as a match, and a proximity-preserving one-way function is used, then the tracking client evaluates whether the re-identification data item matches by testing (3'). The set G P1 can be filled iteratively according to the following rule: such that if the set contains at least one re-identification data item g j,1 such that d(g j,1 , g k,1 ) < D′0, then the new re-identification data item g k+1,1 should be added to the set G P1 .
[0088] In step 312, for each of the sets of mutually matching re-identification data items, the tracking client tracks the corresponding tracking targets P1, P2, P3 based on the positions used to annotate the mutually matching re-identification data items. The output of step 312 can have the format of a "trajectory", that is, the trajectories of the tracking targets P1, P2, P3, from which the position of the tracking target as a function of time can be understood, for example, a table that maps time points to positions and vice versa. Various graphical output formats of step 312 are also possible.
[0089] This concludes the description of the basic functions of the visual tracking system 100 according to the basic architecture. Some alternative embodiments will now be described.
[0090] Optional implementation
[0091] Regarding the architecture of the visual tracking system 100, Figure 1B shows a structure that is different in the following aspects from Figure 1A a different structure.
[0092] The anonymization layer 130 is implemented in a distributed manner. The processing suitable for the anonymization layer 130 is performed in the corresponding processing circuits 121.1, 121.3, 121.4 that are collocated with the corresponding (physical) image sources. This corresponds to Figure 4 process 403, process 404 and Figure 2 steps 212, step 213, step 214 in
[0093] The process of providing feature vectors is also performed in the corresponding processing circuits 121.1, 121.3, 121.4 that are collocated with the image sources. This corresponds to Figure 4 process 401, process 402 and Figure 2 steps 210, step 211 in
[0094] In the common process executed on the processing circuit 121.1, re-identification data items from two (physical) image sources 120.1, image source 120.2 are provided using the first tracking permission token σ1. It can be considered that the image sources 120.1, image source 120.2 form subgroup 111.1. The FOV parts of the image sources 120.1, image source 120.2 overlap.
[0095] Another process executed on processing circuit 121.3 provides a re-identification data item from further image source 120.3 using first tracking authority token σ1. First tracking authority token σ1 is thus specific to the image source group consisting of image source 120.1, image source 120.2, image source 120.3 (and possibly more).
[0096] The two image sources 120.4, 120.5 correspond to different halves of the FOV of a single physical image source. The segmentation of the FOV of the physical image source may be achieved by optical means or digital image processing. In a common process executed on the processing circuit 121.4, the re-identification data items from the two (virtual) image sources 120.4, 120.5 are provided using a second tracking authority token σ2.
[0097] Regarding re-identification / tracking permissions, the relationship between tracking client 140 and image source 120 may be one-to-one, one-to-many, many-to-one, or many-to-many. Here, tracking client 140.1 and tracking client 140.2 are both authorized to be in the FOV of image source 120.1, image source 120.2, and image source 120.3 and use the first re-identification data item g i,1 The second tracking client 140.2 is additionally authorized to be in the FOV of the fourth image source 120.4 and the fifth image source 120.5 and uses the second re-identification data item g i,2 Perform re-identification.
[0098] Figure 1A and Figure 1B The above visible differences between illustrate the architectural variations that can be practiced when implementing the teachings of the present disclosure. These variations can be practiced alone or in various combinations.
[0099] In some embodiments, the scope of visual tracking is defined with respect to time, i.e., by adding step 216 in method 200, in which the anonymization layer 130 stops using the first tracking authority token σ1 after the validity period of the first tracking authority token expires. The anonymization layer 130 may thereafter anonymize the feature vector using a predefined one-way function h modified by a replacement first tracking authority token σ′1 different from σ1. As explained above, the first tracking client will not be able to find a match between the re-identification data items generated before and after the token replacement, even if it considers two re-identification data items generated based on the same feature vector.
[0100] The validity period of the first tracking authority token σ1 may be predetermined, such as every full hour or every day. If not, the expiration time of the validity period may be determined dynamically. To this end, the inventors have envisioned two different but equivalent solutions, which are suitable depending on whether the anonymization layer 130 is implemented using a centralized or distributed architecture.
[0101] In the case of a distributed implementation as shown in Figure 1B , method 200 includes step 215a of negotiating between a first device 130.1 (e.g., processing circuitry 121.1) executing an instance of method 200 using a first tracking entitlement token σ1 and at least one other device 130.2 (e.g., processing circuitry 121.3) executing a corresponding further instance of method 200 using the same first tracking entitlement token σ1. The negotiation step 215a may start with a proposed expiration time expressed in a common time base (network time) from one device, after which the other device(s) unanimously approve the proposal, or at least one other device rejects the proposal when making a counter-proposal for the expiration time. These devices are configured to comply with the approved expiration time.
[0102] Whenever the coordination entity 150 is available in the visual tracking system 100 - that is, regardless of whether the implementation of the anonymization layer 130 is centralized or distributed - the expiration of the validity period of the first tracking entitlement token σ1 can be determined by a decision made by the coordination entity 150. In a distributed implementation ( Figure 1B ), the decision regarding the valid time is sent from the coordination entity 150 to the respective clusters of processing circuitry 121 co-located with the image source 120, which together act as the anonymization layer 130 of the visual tracking system 100. The clusters of processing circuitry that execute the corresponding instances of method 200 and receive the decision of the coordination entity 150 regarding the valid time (step 215b) are configured to behave in accordance with the determined expiration time, i.e., stop using the tracking entitlement token. In a centralized implementation ( Figure 1A ), the coordination entity 150 receives its decision internally, i.e., it makes a decision and behaves accordingly when the expiration time is reached. For the avoidance of doubt, it is clarified that the coordination entity 150 may be responsible for decisions regarding the expiration of the validity period even if the processes related to the anonymization layer 130 are delegated to other components of the visual tracking network 100.
[0103] The expiration of the validity period of the tracking entitlement token mainly affects the anonymization layer. However, it may also be communicated to the tracking clients that use the re-identification data items generated by the anonymization layer, such that these tracking clients can limit their searches (step 311 in method 300) to match the re-identification data items with the validity period of each tracking entitlement token in a timely manner. This can avoid spending processing resources on meaningless searches, as it is known that the tracking clients will never find a match between the re-identification data items generated before and after the replacement of the tracking entitlement token.
[0104] Conclusion
[0105] In summary, a visual tracking system has been proposed in which an image source performs object detection on a captured image to determine the feature vector and location information (e.g., bounding box) of each detected object or person (hereinafter: target). Each image source irreversibly anonymizes the determined feature vector together with a tracking permission token to create an anonymized feature vector (re-identification data item). The image source then sends the re-identification data item together with the associated location information as metadata to a tracking client that is authorized to perform target tracking in the relevant area.
[0106] The proposed solution can be described as a method for tracking targets detected by an image source in a visual tracking system (camera system). Broadly speaking, the following actions are performed.
[0107] Each image source is provided with a set of tracking permission tokens. Each tracking permission token is valid for a predetermined period of time,
[0108] and it implements different rights of the tracking client to track the target at the technical level.
[0109] The image source creates a metadata structure for each target it detects, including the anonymized feature vector and the location information of the detected target. Further, the anonymized feature vector is an anonymization of the combination of the feature vector of the detected target and the corresponding tracking permission token from the set. The anonymization of the combination of the feature vector and the corresponding correct identifier of the set is ideally collision-resistant and irreversible.
[0110] For each image source and each detected target, the metadata structure is transmitted to a tracking client device authorized to track the detected target.
[0111] Each tracking client that receives two or more metadata structures compares the anonymized feature vectors of the different metadata structures. When it finds a match between the anonymized feature vectors, it adds the location information of the two received metadata structures to the trajectory associated with the anonymized feature vector (which in turn can be associated with the detected target).
[0112] The aspects of the present disclosure have been mainly described with reference to several embodiments. However, as will be readily appreciated by those skilled in the art, other embodiments beyond the above-disclosed embodiments are equally possible within the scope of the invention as defined by the appended patent claims.
Claims
1. A visual tracking system, comprising: A plurality of image sources are configured to provide images M within a corresponding field of view, each image source being configured to: detect a sub-region m containing a tracking target in the image obtained from the image source i ; and calculating, for each sub-region, a feature vector f(m i ); An anonymization layer, configured to: anonymize each feature vector using a predefined one-way function h modified by the first tracking authority token σ1 to provide a first re-identification data item g i,1 =h([f(m i ), σ1]), the first tracking authority token is specific to a first set of at least one image source in the image sources; and a second re-identification data item g is provided by anonymizing each feature vector using a predefined one-way function h modified by a second tracking authority token σ2 i,2 =h([f(m i ), σ2]), a second tracking permission token is specific to a second set of at least one image source among the image sources and is different from the first tracking permission token; And publicly annotate the location X(m) of the corresponding sub-area to the tracking client i ), while preventing access to the feature vector; A plurality of tracking clients are configured to perform re-identification on a tracking target using data obtained from the anonymization layer, including: a first tracking client, the first tracking client being authorized to perform re-identification in a field of view of a first group of the image sources and obtain the first re-identification data item with a location annotation; and a second tracking client, the second tracking client being authorized to perform re-identification in a field of view of a second group of the image sources and obtain the second re-identification data item with a location annotation, Wherein, the anonymization layer is implemented in a processing circuit separate from the tracking client.
2. The visual tracking system according to claim 1, wherein: The anonymization layer is implemented in processing circuitry of a coordinating entity in the visual tracking system and / or in processing circuitry co-located with a respective one of the image sources.
3. A method for providing anonymized data to facilitate re-identification in a visual tracking system, the method comprising: In the images M obtained from multiple image sources, sub-regions m each containing a tracking target are detected. i ; Calculating, for each sub-region, a feature vector representing the visual appearance of the tracking target therein; For the first subset of image sources: providing a first re-identification data item g by anonymizing each feature vector using a predefined one-way function h modified by a first tracking authority token σ1 i,1 =h([f(m i ), σ1]); and publicly annotating the position X(m) of the corresponding sub-area to the first tracking client. i ) of the first-re-identification data item; For the second subset of the image sources: providing a second re-identification data item g by anonymizing each feature vector using a predefined one-way function h modified by a second first tracking authority token σ different from the first tracking authority token i,2 =h([f(m i ), σ2]); and publicly disclose the second-re-identification data item annotated with the location of the corresponding sub-area to a second tracking client; Prevents access to the feature vector.
4. The method according to claim 3, wherein: The first tracking authority token σ1 is specific to the image source group including the first sub-group, and / or the second tracking authority token σ2 is specific to the image source group including the second sub-group.
5. The method according to claim 3, further comprising: After the validity period of the first tracking authority token σ1 expires, the first tracking authority token σ1 is no longer used.
6. The method according to claim 5, further comprising: Negotiation regarding expiration of the validity period of the first tracking authority token σ1 is performed between the device executing the method and at least one other device using the first tracking authority token σ1.
7. The method according to claim 5, further comprising: An indication of the expiration of the validity period of the first tracking authority token σ1 is received by a coordinating entity in a visual tracking system.
8. The method according to claim 5, wherein: After the expiration of the validity period of the first tracking authority token σ1 , the feature vectors of the first subset of the image sources are anonymized using the predefined one-way function h modified by the replaced first tracking authority token σ′1 .
9. The method according to claim 3, wherein: The feature vector f(m i ) is calculated to take a value in a discrete set; and / or Used to provide the re-identification data item g i,1 , g i,2 The one-way function h is proximity preserving.
10. The method according to claim 3, wherein: Said provision of said re-identification data item comprises applying as cryptographic salt a tracking authority token σ1 , σ2 modifying said predefined one-way function.
11. The method according to claim 3, wherein: The tracking targets include humans and / or inanimate objects.
12. A method for visually tracking a tracking target in a field of view of a first group of image sources, the method comprising: Receive the first re-identification data item g i,1 , each first-recognition data item is derived from the feature vector f(m i ), the feature vector represents the visual appearance of the tracking target in a sub-region of an image obtained from an image source in the first group, and each first re-identification data item is annotated with a location of the corresponding sub-region, wherein the first re-identification data item has been calculated using a predefined one-way function h modified by a first tracking authority token σ1 of the first group specific to the image source; searching the first re-identification data items for matching re-identification data items; and For a set of mutually matching re-identification data items, a tracking target is tracked based on the positions used to mark the mutually matching re-identification data items.
13. The method according to claim 12, wherein: The search for matching re-identified data items is limited in time to the validity period of the first tracking authority token σ1.
14. A device or a cluster of devices comprising a processing circuit configured to perform the method of claim 3.
15. A computer program comprising instructions which, when executed by a computer, cause the computer to perform the method according to claim 3.
Citation Information
Patent Citations
Evaluation device for re-identification and corresponding method, system and computer program
EP3971734A1
Image anonymization apparatus, image anonymization method, and program
JP2021064089A
Image anonymization apparatus, image anonymization method, and program
JP2021064203A
Using object re-identification in video surveillance
US20180374233A1
Privacy-preserving demographics identification
US20190034716A1