Method for the management of data associated with multimedia contents

The method addresses privacy concerns in facial recognition by creating a composite image of faces from multimedia contents, enabling efficient searches while ensuring user consent and minimizing data storage risks.

WO2025109505A1PCT designated stage expired Publication Date: 2025-05-30PICA GRP SPA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/061658
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-24
Filing Date
2024-11-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing facial recognition systems in multimedia contents face challenges in managing and recognizing faces while ensuring user privacy, particularly in cases where explicit consent for data storage is not obtained.

Method used

A method that collects faces from multiple multimedia contents into a composite image, allowing for efficient search and management of faces without the need to store individual face data, while implementing privacy controls through consent-based inclusion in the composite image.

Benefits of technology

This approach optimizes resource usage for face searches, enhances user privacy by minimizing data storage and unauthorized access, and allows for consent-based management of facial data, thereby complying with privacy regulations like GDPR.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024061658_30052025_PF_FP_ABST
    Figure IB2024061658_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention refers to a method (100) for the management of faces extracted from multimedia contents, the method comprising the steps of detecting (S110) one or more first faces in a first multimedia content, where the first multimedia content is identifiable by means of a first piece of identification data, detecting (S111) one or more second faces in a second multimedia content, where the second multimedia content is identifiable by means of a second piece of identification data, generating (S120) a composite image comprising - the one or more first faces, - the first piece of identification data associated with the one or more first faces, - the one or more second faces - the second piece of identification data associated with the one or more second faces.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR THE MANAGEMENT OF DATA ASSOCIATED WITH MULTIMEDIA CONTENTS

[0002] DESCRIPTION

[0003] The present invention concerns the field of facial recognition in multimedia contents and describes in particular a method for the management of faces extracted from multimedia contents , and subsequent recognition of any face from among the latter, which improves the privacy of the persons shown in the multimedia contents .

[0004] STATE OF THE ART

[0005] Facial recognition systems are generally based on a numerical comparison, understood as calculation of a distance in a Euclidean or non-Euclidean geometrical space between a vector of N points associated with the face of a person, and other M vectors or N points , each of which represents another face present in a generic database .

[0006] In general , the matrix or vector representation of a face is defined as the determination and sampling of some points corresponding to somatic traits , and i f necessary the calculation of generic distances between the points analysed, which redefine in matrix representation the face considered for the calculation . There are many techniques and many definitions of how these points should be extracted and which distances should be calculated; also the number of points extracted from a face is variable and each recognition system defines its own characteristics .

[0007] A common factor of the various systems , regardless of the method used for calculation of the data, is the search procedure described above , which by definition entails storage of the comparison faces in an organi zed data structure such as , for example , a database . Said storage has implications in terms of privacy of the user, whose face is saved in this data structure , in particular in compliance with the GDPR, since the vector representation of the face constitutes a piece of biometric data . Possible unauthori zed access to a database that stores the vector representations of the faces of generic persons would allow, for example , comparison of said data with faces obtained from photographs or videos in other situations and contexts , with implications in terms of unauthori zed tracking of persons .

[0008] For example , it would be possible to find a photograph of a person from a social media channel , calculate the vector representation of his / her face , and compare it with a generic database of vector data, and see whether said person is present or not therein, thus ascertaining whether the person in question has , for example , participated in a given event or whether he / she is present in the context relative to the vector data stored in the above-mentioned database .

[0009] Another example is a system that allows generic photographers to upload on an online platform photographs taken in a context , and allows a user to access the photographs or videos that show him / her by searching for his / her face from among the photographs uploaded . I f the consent of the participants in a certain event to which the photographs refer has not been obtained beforehand, there would be the risk of storing the participant ' s face in a manner not explicitly authori zed by the participant , which is not permitted by the GDPR .

[0010] Even i f photographers take photographs of users who have consented to storage of the piece of biometric data associated with their face , further complications arise i f third persons , for example , appear in said photograph such as , for example , the public at an event , who have no way of explicitly expressing their consent to the storage of their face .

[0011] It is therefore necessary to solve the problems arising from storage on a data structure o f the representation of faces for which no explicit consent has been obtained . SUMMARY OF THE INVENTION

[0012] The invention is defined by the independent claims . The dependent claims define advantageous embodiments of some aspects of the invention .

[0013] The invention is based generally on the concept that faces present in a plurality of multimedia contents can be collected in a composite image , thus facilitating a subsequent search for one or more faces in said multimedia contents , by searching in the composite image .

[0014] This approach is particularly advantageous for the reasons that will become evident from the following description, and in general because it allows optimi zation of the resources necessary for the search and improved management of the privacy of the persons shown in the multimedia contents .

[0015] BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 illustrates schematically a method 100 for the management of faces extracted from a multimedia content . Figure 2 illustrates schematically some steps of a method 200 for the management of faces extracted from a multimedia content .

[0017] DESCRIPTION OF PREFERRED EMBODIMENTS

[0018] Figure 1 illustrates schematically a method 100 for the management of faces extracted from a multimedia content . As can be seen in figure 1 , the method 100 for the management of faces extracted from multimedia contents comprises a detection step S 110 for detecting one or more first faces in a first multimedia content. The step Sill can be implemented with any algorithm, known per se, for detecting faces in a multimedia content. The multimedia contents in the context of the present invention can be images or videos.

[0019] In the invention, the first multimedia content is identifiable by means of a first piece of identification data. The first piece of identification data is a piece of data having a one-to-one correspondence with the first multimedia content, which allows identification of the first multimedia content from among the plurality of multimedia content. It will be evident to persons skilled in the art that this can be implemented in different ways, for example by means of a hash of the multimedia content, by means of a unique name assigned to the multimedia content, etc.

[0020] The method 100 further comprises a detection step Sill for detecting one or more second faces in a second multimedia content, where the second multimedia content is identifiable by means of a second piece of identification data. The second multimedia content is different from the first multimedia content. The considerations already made for step S110 apply also to step Sill. Analogously to what has already been described for the first piece of identification data, the second piece of identification data is a piece of data having one-to-one correspondence with the second multimedia content.

[0021] It will be clear to persons skilled in the art that a number N of similar detection steps applicable to N multimedia contents can be implemented, and that the detection steps and the multimedia contents are therefore not necessarily limited to two. This is shown schematically by step SUN in figure 1, which can therefore be executed successively for an indefinite number of repetitions, associated with as many multimedia contents.

[0022] The result of the detection steps S111-S11N is therefore the identification of a series of faces, in a plurality of multimedia contents. By way of example, the following situation is possible:

[0023] Face #1, detected in the first multimedia content identified by the piece of identification data ID#1 Face #2, detected in the first multimedia content identified by the piece of identification data ID#1 Face #3, detected in the second multimedia content identified by the piece of identification data ID#2 Face #4, detected in the second multimedia content identified by the piece of identification data ID#2 Face #5, detected in the third multimedia content identified by the piece of identification data ID#3 Face #6, detected in the fourth multimedia content identified by the piece of identification data ID#4

[0024] The method 100 lastly comprises a generation step S120 for generating a composite image comprising the one or more first faces, identified in the first multimedia content, where the first piece of identification data is associated with the one or more first faces, the one or more second faces, identified in second multimedia content, where the second piece of identification data is associated with the one or more second faces.

[0025] It will be evident that if further detection steps SUN have been executed, the corresponding results can be incorporated in the composite image in an analogous manner. The successive steps illustrated in figure 1 are to be understood as optional and will be described below .

[0026] The composite image is there fore a digital image , comprising the plurality of faces , extracted from the respective multimedia contents . It will be clear to persons skilled in the art that , according to the detection algorithm used, it will be possible to extract only the face , or the face and a region of the multimedia content in the immediate vicinity thereof . By way of example , algorithms are known that allow the recognition of a rectangular area of a multimedia content comprising a face . In the context of the composite image , the reference to the face can be interpreted as the region of the multimedia contents which the recognition algorithm employed in step S 110-S 11N has identi fied as comprising a face .

[0027] In addition, the composite image comprises , or is associated with, information that associates the piece of identi fication data of the multimedia content with the faces that have been identi fied therein .

[0028] The association can, for example , be saved in the metadata of the files of the digital image . Alternatively, a digital file comprising the association can be saved separately from the file of the digital image and be associated uniquely with the digital image .

[0029] The association can be implemented in di f ferent ways . For example , the faces in the digital image can be numbered, and the association can be implemented by saving the number of a given face in association with the piece of identi fication data of the respective multimedia content . Alternatively, or additionally, the faces in the digital image can be positioned at di f ferent coordinates in space of the image , and the association can be implemented by saving the coordinates of a given face in association with the piece of identi fication data of the respective multimedia content .

[0030] Thanks to this implementation of the method 100 it is therefore possible to obtain a composite image comprising the faces detected in one or more multimedia contents . As will be evident from the following description, this allows a successive search for faces in the composite image , and relative determination of the respective multimedia contents , without requiring identi fication of the faces of all the multimedia contents , whenever a new face has to be searched for .

[0031] Furthermore , in some embodiments , the method 100 can further comprise a step, not illustrated, that excludes predetermined faces from inclusion in the composite image . This step consists , for example , in excluding faces that correspond to users that have not given their consent to the processing o f their personal data . In this way it is possible for a user to load his / her face and indicate that he / she does not wish to be identi fied in the multimedia content . The respective face can be saved in a database of excluded faces , and in the step described here , not illustrated, it will be possible to recogni ze whether the face detected in the multimedia contents is present in said database and, i f present , include it in the generation of the composite image carried out in step S 120 .

[0032] Alternatively, or additionally, in some embodiments , the method 100 can further comprise a step, not illustrated, which, in step S 120 , allows only predetermined faces to be included in the composite image . This step consists , for example , in including only faces that correspond to users who have given their consent to processing of their personal data . In this way it is possible for a user to load his / her face and indicate that he / she wishes to be identi fied in the multimedia contents . The respective face can be saved in a database of faces not excluded, and in the step described here , not illustrated, it will be possible to recogni ze whether the face detected in the multimedia contents is present in said database , and only in this case include it in generation of the composite image carried out in step S 120 .

[0033] In this way it i s possible to implement in a relatively simple manner the management of user privacy, controlling which faces are included in the composite image . Advantageously, this is possible without modi fying the original multimedia contents .

[0034] It is therefore evident that , via implementation of the steps described so far, it i s possible to create a composite image , comprising the faces of people recogni zed in a plurality of multimedia contents . The grouping of said faces in the composite image advantageously facilitates the subsequent searches and also the implementation of a privacy control .

[0035] In some embodiments , as can be seen in figure 1 , the method 100 further comprises a step S 130 of obtaining a face to be searched for . The step S 130 can be implemented by a user, for example , as upload of a sel fie or an image comprising the face to be searched for .

[0036] The method 100 can furthermore comprise a step S 140 of calculating at least one biometric value of the face to be searched for . The calculation of the at least one biometric value can be implemented in a manner known per se , by means of an algorithm that associates at least one biometric value with the face . In general , said algorithm associates a vector of N points with the user' s face , giving the N points of the vector numerical values which are calculated based on biometric characteristics of the face . At a step S141, which is to be considered optional, the face to be searched for can be eliminated, preferably immediately after the step S140. This is particularly advantageous since it allows the storage time of the user' s face to be minimized, said storage entailing various obligations and risks in terms of privacy. Furthermore, the reduced duration of the storage time advantageously means that storage on long-term supports, such as harddisks, can be avoided and storage can be implemented exclusively on short-term supports, such as RAM memory, and more precisely in an algorithm stack in a protected and inviolable memory space, which are known to be more difficult to access in the event of hacking of the system that implements the method 100, advantageously reducing the risks of access to personal information also in the case of hacker attacks.

[0037] The method 100 can furthermore comprise a recognition step S160, based on the biometric value, for recognizing the face to be searched for from among the faces of the composite image, and an identification step S170 for identifying one or more multimedia contents based on the respective piece of data corresponding to the one or more faces identified in the recognition step S160.

[0038] The recognition step S160 can be implemented with an algorithm, known per se, which, based on the biometric value, allows the recognition of a corresponding face, if present, from the faces of the composite image. Once a corresponding face has been identified, the step S170 can be implemented recovering the piece of identification data associated, or the identification data associated if the face is present in several multimedia contents, with the face identified. This allows identification of the multimedia content, or multimedia contents, in which the face is present , requiring only the recognition thereof in the composite image, without requiring access to the multimedia contents .

[0039] This approach is particularly advantageous for various reasons . The face to be searched for is searched for in the faces present in the composite image , instead of in the starting multimedia contents , speeding up the search operation and requiring far fewer resources . Furthermore , access to the multimedia contents is limited, thus reducing the risk of hacker attacks . In addition, the face search algorithm that operates at step S 160 has access only to the faces that are present in the composite image which, as described above , can be advantageously fewer than the faces present in the multimedia contents following a first filtering step of the faces on the basis of privacy indications by the users . The biometric data of the faces in the composite image can furthermore be created at the time of the search carried out in step S 160 and, additionally or alternatively, can be immediately eliminated after step S 160 , avoiding long-term saving of the respective biometric data .

[0040] In some embodiments , the method 100 can further comprise a step, not illustrated, of eliminating the face to be searched for immediately after the calculation step S 150 . In this way it is possible to obtain an increase in privacy and a reduction in the risk of access to the face by hackers , thanks to reduction in the storage time of the face in a clear manner . The considerations relating to step S 141 apply also to the step j ust described .

[0041] In the preceding description, the composite image has been described as comprising the faces relevant to the steps S 110-S 11N . In some embodiments , the detection step S 110 of the one or more first faces in the first multimedia content , or analogously the steps S i l l and S UN, can further comprise the determination of identi fication coordinates of a region containing each of the one or more faces detected . The region containing the face , known also as bounding box in face recognition algorithms , is a geometric form, generally a rectangle , within which the recogni zed face is inscribed .

[0042] In this case , the generation step S 120 of the composite image can further comprise the association of the one or more first faces with the respective coordinates of the region comprising the corresponding one or more first faces detected in the first multimedia content , the association of the one or more second faces with the respective coordinates of the region comprising the corresponding one or more second faces detected in the second multimedia content , the association of the one or more N faces with the respective coordinates of the region comprising the corresponding one or more N faces detected in the Nthmultimedia content .

[0043] In this way, together with the identi fication of the face in the starting multimedia content , it is also possible to identi fy the region of the multimedia content that contains the face and save the respective coordinates thereo f in a manner associated with the face . The association of the coordinates of this region with the respective face can be carried out analogously to what has already been described for the association between the face and the piece of identi fication data of the multimedia content .

[0044] This allows , for example , the user carrying out the step S 130 to be provided with a cutout of the multimedia content in a simpli fied manner ; in particular the user can be provided with the region associated with the face in the multimedia content . In other words , once the face of the user is recogni zed in step S 160 , it is possible to obtain the identi fied coordinates of the region of the multimedia content that corresponds to the face , and the content of said region can be provided to the user, for example to confirm that the search has been carried out correctly and that the face indicated in the region actually corresponds to the face provided in step S 130 .

[0045] Alternatively, or additionally, by means of a step not illustrated, it is possible to anonymi ze the users visible in the multimedia content outside the coordinates of the face detected . Algorithms for anonymi zing faces in multimedia contents are known per se to persons skilled in the art . Knowing the coordinates of the region in which the face resulting from the search is positioned allows these algorithms to be executed on the remaining part of the multimedia content , thus guaranteeing the anonymity of other persons visible in the multimedia content , maintaining the face searched for unchanged .

[0046] Figure 2 schematically illustrates some steps of a method 200 for the management of faces extracted from a multimedia content . The method 200 di f fers from the method 100 due to the addition of some steps , which will be described below . In particular, the method 200 comprises a detection step

[0047] 5210 for detecting one or more first bodies of the persons corresponding to the one or more first faces in the first multimedia content , and, analogously, a detection step

[0048] 5211 for detecting one or more second bodies of the persons corresponding to the one or more second faces in the second multimedia content . Also in this case , as already discussed for steps S 110-S 11N, it will be possible to implement steps S210 , S211 N times , as schematically illustrated by step 21N .

[0049] Thanks to recognition of the bodies of the persons corresponding to the faces already detected, it is possible to recogni ze any visual markers on the bodies and use these markers to decide whether to include or not the respective faces in the composite image .

[0050] In particular, as can be seen in figure 2 , a possible alternative implementation of the generation step S 120 can be implemented from the generation step S220 , which comprises a determination step S221 for determining the presence of a predetermined marker on the body relative to a given face . I f the marker is identi fied, the face is discarded in step S223 and the generation step S220 returns to step S221 , applied to the body associated with the next face . I f the marker is absent , the generation step S220 proceeds to step S222 , in which the composite image is generated as previously described for figure 1 .

[0051] Repeating the generation step D220 for all the faces detected in steps S 110-S 11N, the composite image will therefore be composed of the one or more f irst faces for which the respective one or more first bodies do not contain a predetermined marker, the first piece of identi fication data associated with the one or more first faces thus chosen, the one or more second faces for which the respective one or more second bodies do not contain a predetermined marker, the second piece of identi fication data associated with the one or more second faces thus chosen, etc . Recognition of the marker on the body therefore allows the respective face to be excluded from the database within which the subsequent search is carried out , allowing event participants to indicate their intention not to be found in the multimedia contents by simply applying a predetermined marker on the body . This application can be particularly advantageous , for example , in the case of events where there are visitors and workers , with the latter recogni zable from markers and not needing, or not wishing, to be traceable in the multimedia contents recorded during the event .

[0052] It will be evident that the contrary implementation is also possible , in which only the faces corresponding to a body on which the marker is present will be included in the composite image .

[0053] In this case , at the generation step S220 the composite image will comprise the one or more f irst faces for which the respective one or more first bodies contain a predetermined marker, the first piece of identi fication data associated with the one or more first faces thus chosen, the one or more second faces for which the respective one or more second bodies contain a predetermined marker, the second piece of identi fication data associated with the one or more second faces thus chosen, etc .

[0054] Recognition of the marker on the body therefore allows the respective face to be included in the database within which the subsequent search is carried out , allowing event participants to indicate their intention to be found in the multimedia contents by simply applying a predetermined marker on the body . This application can be particularly advantageous , for example , in the case of events where there are participants who wish to be visible in the contents , and spectators who wish to remain anonymous , with the former being recogni zable by markers and wishing to be traceable in the multimedia contents recorded during the event .

[0055] It is evident that the two approaches can be furthermore combined, using a first marker to indicate the wish to be excluded from the composite image , and a second marker indicating the wish to be included, with the first marker di f ferent from the second marker .

[0056] Although the invention has been described so far in terms of a method, it i s clear that implementations in the form of devices are also possible . In particular, it will be possible to implement a device comprising a processor and a memory, by way of example a PC, a smartphone , a server, or similar , where the memory comprises instructions configured to cause in the processor the execution of one or more steps of any one of the methods previously described .

[0057] In addition, it will be evident to persons skilled in the art that , although speci fic embodiments of the invention have been described and / or illustrated as comprising a plurality of characteristics , each of these characteristics can be implemented together with any other of these characteristics , without all the other characteristics described or illustrated having to be implemented in the combination .

Claims

CLAIMS1. Method (100, 200) for the management of faces extracted from multimedia contents, the method comprising the steps of detecting (S110) one or more first faces in a first multimedia content, where the first multimedia content is identifiable by means of a first piece of identification data, detecting (Sill) one or more second faces in a second multimedia content, where the second multimedia content is identifiable by means of a second piece of identification data, generating (S120, S220) a composite image comprising the one or more first faces, where the first piece of identification data is associated with the one or more first faces, the one or more second faces, where the second piece of identification data is associated with the one or more second faces.

2. The method (100, 200) according to claim 1, also comprising the steps of obtaining (S130) a face to be searched for, calculating (S140) at least one biometric value of the face to be searched for, recognizing (S160) , on the basis of the biometric value, the face to be searched for among the faces of the composite image, identifying (S170) one or more multimedia contents on the basis of the respective piece of data corresponding with the one or more faces identified in the recognizing step (S160) .

3. The method (100, 200) according to claim 2, also comprising the step ofeliminating the face to be searched for immediately after the calculating step (S150) .

4. The method (100, 200) according to any one of the preceding claims, where the first piece of identification data is a piece of data in one-to-one correspondence with the first multimedia content, and / or the second piece of identification data is a piece of data in one-to-one correspondence with the second multimedia content .

5. The method (100, 200) according to any one of the preceding claims, where the step of detecting (S110) the one or more first faces in the first multimedia content also comprises the determination of identification coordinates of a region containing each of the one or more first faces, the step of detecting (Sill) the one or more second faces in the second multimedia content also comprises the determination of identification coordinates of a region containing each of the one or more second faces, the step of generating (S120, S220) the composite image also comprises the association of the one or more first faces with the respective coordinates, the association of the one or more second faces with the respective coordinates.

6. The method (200) according to any one of the preceding claims, also comprising the steps of detecting (S210) one or more first bodies of the people corresponding with the one or more first faces of the multimedia content,detecting (S211) one or more second bodies of the people corresponding with the one or more second faces of the multimedia content, and where, in the generating step (S220) , the composite image comprises the one or more first faces for which the respective one or more first bodies do not contain a predetermined marker, the first piece of identification data associated with the one or more first faces, the one or more second faces for which the respective one or more second bodies do not contain a predetermined marker, the second piece of identification data associated with the one or more second faces.

7. The method (200) according to any one of the preceding claims, also comprising the steps of detecting (S210) one or more first bodies of the people corresponding with the one or more first faces of the multimedia content, detecting (S211) one or more first bodies of the people corresponding with the one or more first faces of the multimedia content, and where, in the generating step (S220) , the composite image comprises the one or more first faces for which the respective one or more first bodies contain a predetermined marker, the first piece of identification data associated with the one or more first faces,the one or more second faces for which the respective one or more second bodies contain a predetermined marker, the second piece of identi fication data associated with the one or more second faces .8 . Device comprising a processor and a memory, where the memory compri ses instructions configured so as to cause in the processor the execution of the method ( 100 , 200 ) according to any one of the preceding claims .

Citation Information

Patent Citations

  • Identification of individuals in images and associated content delivery

    US20170255820A1