Method and device for video communication
The method and device for video communication address the issue of unauthorized individuals in video frames by using image processing to detect and remove them, ensuring privacy and professionalism in video communications.
Patent Information
- Application Number
- FR2023014754
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-27
AI Technical Summary
Existing video communication systems struggle to effectively manage situations where unauthorized individuals inadvertently enter the video frame, leading to privacy concerns and unsuitable background exposure.
A method and device for video communication that utilize image processing to detect and verify individuals in the frame, erase unauthorized persons from the image, and transmit the processed image, incorporating face and body detection, association, and classification using a database of authorized individuals.
This solution ensures that only authorized individuals remain visible in the video communication, enhancing privacy and maintaining a professional environment by effectively removing unauthorized persons from the video frame.
Smart Images

Figure 00000021_0000 
Figure 00000023_0000 
Figure 00000026_0000
Abstract
Description
Title of the invention: Method and device for video communication Technical field
[0001] A method and device for video communication are described. The method and device may be used, for example, in a video calling or videoconferencing application. Technical background
[0002] Video call or videoconferencing systems have found numerous applications, both in the professional and private fields, or across both fields, particularly through teleworking. The boundary between the private sphere and the professional environment has thus become permeable. Thus, a video call can be considered an intrusion, due to the information it provides to a user on the physical, family or professional environment of his interlocutors. Various solutions have been proposed, including the possibility of blurring the background of the image, or even superimposing a virtual background on it. However, such solutions are not suitable for situations where people make an untimely intrusion into the plane of the image. Summary
[0003] One or more embodiments relate to a method of video communication implemented by a device comprising a processor and a memory comprising software code, the processor executing the software code causing the device to implement the method, the method comprising:
[0004] - obtaining an image generated by a camera;
[0005] - the detection of one or more persons in the image;
[0006] - in case of detection of one or more persons, the verification for each person of a criterion indicating whether the person should or should not be part of the video communication;
[0007] - processing the image to erase the detected person(s) from the image not to be part of the video communication;
[0008] - transmission of the processed image.
[0009] According to one or more embodiments, the detection comprises identifying one or more first areas of the image, each first area comprising a face.
[0010] According to one or more embodiments, the detection further comprises:
[0011] - the identification of one or more second zones of the image, each second zone comprising a body;
[0012] - the association, according to an association criterion, of a first zone and a second zone to form a representation of a person in the image.
[0013] According to one or more embodiments, the method comprises, for a given first area that cannot be associated with a second area based on the association criterion, determining a third area, where the third area is an area of the image that is a function of the given first area and is intended to serve as a second area associated with the given first area to form a representation of a person in the image for image processing.
[0014] According to one or more embodiments, the method comprises, when a second zone cannot be associated with a first zone, marking this second zone as part of a person who should not participate in the video communication, the representation of this person then comprising only the second zone.
[0015] According to one or more embodiments, the method comprises extracting, from each first zone, characteristic parameters of the face of each first zone, said characteristic data being adapted to make it possible to determine, from a database, whether a person corresponding to a face should or should not be part of the video communication.
[0016] According to one or more embodiments, said database comprises:
[0017] either characteristic parameters of faces for one or more persons authorized to participate in the video communication;
[0018] or at least one of:
[0019] - characteristic parameters for one or more faces as well as for each face an indication that the person corresponding to the face should be part of the video communication; and
[0020] - characteristic parameters for one or more faces as well as for each face an indication that the person corresponding to the face should not be part of the video communication.
[0021] According to one or more embodiments, the method comprises initializing the database with data stored in advance.
[0022] According to one or more embodiments, the method comprises augmenting the database by:
[0023] - the identification of one or more first zones comprising a face, during a time interval of the start of the video communication;
[0024] - the extraction, from each first zone, of characteristic parameters of the face of each first zone;
[0025] - storing the characteristic parameters of the face of each first zone with a respective indication that the person corresponding to the face should be part of the video communication.
[0026] According to one or more embodiments, the processing of the image comprises
[0027] - obtaining a mask from the representations of people who should not be part of video communication;
[0028] - obtaining a processed image according to the mask, of the image obtained from from the camera and a previously obtained processed image.
[0029] According to one or more embodiments, the method comprises, following the processing of the image to erase from the image the detected person(s) not to be part of the video communication, a cropping of the processed image to obtain a cropped processed image, the cropping being configured to remove at least left and right side bands from the processed image, the side bands being of a width capable of containing one or more persons before their detection becomes possible, the transmission being carried out with the cropped processed image.
[0030] Another aspect relates to a device comprising a processor and a memory comprising software code, the processor executing the software code causing the device to implement one of the methods described and in particular one of the methods above.
[0031] Another aspect relates to a television decoder comprising a device such as above.
[0032] Another aspect relates to a computer program product comprising instructions which when executed by at least one processor cause the execution of one of the methods described and in particular one of the methods above.
[0033] Another aspect relates to a non-transitory storage medium comprising instructions which when executed by at least one processor cause the execution of one of the methods described and in particular one of the methods above. Brief description of the figures
[0034] Other characteristics and advantages will become apparent when reading the detailed description which follows, for the understanding of which reference will be made to the attached drawings among which:
[0035] [Fig-1] - [Fig.l] is a schematic diagram of a system comprising a device according to one or more embodiments.
[0036] [Fig.2] - [Fig.2] is a flowchart of a method according to one or more exemplary embodiments;
[0037] [Fig.3] - [Fig.3] is a schematic diagram of an image before its processing and after its processing by a method according to one or more embodiments;
[0038] [Fig.4] - [Fig.4] is a flowchart of the method according to one or more exemplary embodiments;
[0039] [Fig.5] - [Fig.5] is a flowchart detailing sub-steps of a step of [Fig.4] according to one or more exemplary embodiments;
[0040] [Fig.6] - [Fig.6] is a schematic diagram illustrating the detection of bodies and faces according to a particular exemplary embodiment;
[0041] [Fig.7] - [Fig.7] is a schematic diagram illustrating the principle of body-face association;
[0042] [Fig.8] - [Fig.8] is a schematic diagram illustrating the obtaining of a fictitious body for an unassociated face;
[0043] [Fig.9] - [Fig.9] is a schematic diagram illustrating the obtaining of two vectors from two respective faces;
[0044] [Fig. 10] - [Fig. 10] is a schematic diagram illustrating the principle of classification according to one or more embodiments;
[0045] [Fig. 11] - [Fig. 11] is a schematic diagram illustrating two variants of realization of a mask of a current image;
[0046] [Fig. 12] - [Fig. 12] is a schematic diagram illustrating the different stages of the construction of a mask implementing semantic segmentation according to an exemplary embodiment;
[0047] [Fig. 13] - [Fig. 13] is a schematic diagram illustrating the composition of the image S_t;
[0048] [Fig. 14] - [Fig. 14] is a schematic diagram illustrating the problem of a person not being detected because they are partially visible in the image captured by the camera;
[0049] [Fig. 15] - [Fig. 15] is a schematic diagram of a method for reducing the impact of an undetected person near an edge of the image;
[0050] [Fig. 16] - [Fig. 16] is a schematic diagram illustrating the alignment of a face according to one or more exemplary embodiments. Detailed description
[0051] In the following description, identical, similar or analogous elements will be designated by the same reference numerals. The block diagrams, algorithms and message sequence diagrams in the figures illustrate the architecture, functionalities and operation of systems, devices, methods and computer program products according to one or more exemplary embodiments. Each block of a block diagram or each phase of an algorithm may represent a module or a portion of software code comprising instructions for implementing one or more functions. According to certain implementations, the order of the blocks or phases may be changed, or the functions corresponding may be implemented in parallel. The process blocks or phases may be implemented using circuits, software, or a combination of circuits and software, in a centralized manner, or in a distributed manner, for all or some of the blocks or phases. The systems, devices, processes, and methods described may be modified, added to, and / or deleted, while remaining within the scope of this description. For example, the components of a device or system may be integrated or separated. Also, the described functions may be implemented using more or fewer components or phases, or with other components or through other phases. Any suitable data processing system may be used for the implementation.A suitable data processing system or device comprises, for example, a combination of software code and circuits, such as a processor, controller or other circuit adapted to execute the software code. When the software code is executed, the processor or controller causes the system or device to implement all or part of the functionalities of the blocks and / or phases of the processes or methods according to the exemplary embodiments. The software code may be stored in non-volatile memory or on a non-volatile storage medium (USB key, memory card or other medium) readable directly or through a suitable interface by the processor or controller.
[0052] [Fig.l] is a schematic diagram of a system illustrating one or more embodiments in a non-limiting manner. The system of [Fig.l] comprises a device 100 and a display screen 101. The device 100 comprises a processor 105, a non-volatile memory 106 comprising software code, and a working memory 107.
[0053] Furthermore, the various components of the device 100 are controlled by the processor 105, for example via an internal bus 110.
[0054] The device 100 may further comprise an interface (not illustrated) by which it is connected to the screen 101. This interface is for example an HDMI interface. The device 100 is adapted to generate a video signal for display on the screen 101. The generation of the video signal is for example carried out by the processor 104. The device 100 further comprises an interface 111 for connecting it to a communication network, for example the internet network.
[0055] The device 100 further comprises a camera 104 and a microphone 112. The software code comprises a video communication application (video call, video conference, etc.) implementing the camera and the microphone.
[0056] The device 100 may optionally be controlled by a user 102, for example using a user interface, shown here in the form of a remote control 103. The device 100 may further optionally comprise an audio source, illustrated by two speakers 108 and 109. The device may optionally- It may also include a hardware neural processing unit (NPU) whose function is to accelerate the calculations required by a neural network.
[0057] In some contexts, the device 100 is for example a digital television receiver / decoder, while the display screen is a television. However, the invention is not limited to this specific context and can be used in the context of any video communication, such as for example a video communication application on a mobile phone, a computer, etc.
[0058] The system of [Fig.l] is given for illustrative purposes for the clarity of presentation of the exemplary embodiments and a current implementation may comprise more or fewer components. Furthermore, certain components described as integrated into the device 100 may be external to this device and connected to it via a suitable interface - this is notably the case for the camera 104 or the microphone 112. Conversely, certain components of the system described as external to the device 100 may be integrated into the device - for example the display screen or the user interface 103.
[0059] [Fig.2] is a flowchart of a method according to one or more exemplary embodiments. In 201, an image is captured by the camera 104. In 202, the image is analyzed to detect whether one or more people are present. In 203, a classification of the detected people is carried out. This classification indicates whether a person must be masked or not. The image is transformed in 204 to mask, if necessary, the people who must be masked.
[0060] The resulting image is then ready for transmission to one or more recipients, at 205. This transmission may be preceded by other processing of the image and / or of the preceding and following images, such as compression, the addition of elements in the image, etc.
[0061] In the following, the term 'unauthorized person' ('NP') will refer to a person to be masked and the term 'authorized person' ('AP') will refer to a person not to be masked.
[0062] [Fig. 3] schematically illustrates an image before its processing (It, top image) and after its processing (St, bottom image). Image It shows a room filmed by the camera. In this room there are two people, one of whom is authorized ('PA') and the other is not authorized ('PN'). In the transformed image, the unauthorized person is masked, that is to say, it is hidden. In the example of [Fig. 3], the person PN is replaced by the background of the room which would appear if the person PN were absent.
[0063] [Fig.4] is a flowchart of the method according to one or more exemplary embodiments. This flowchart is more detailed than that of [Fig.2].
[0064] In 401, a database of authorized persons is initialized.
[0065] According to one or more exemplary embodiments, the database comprises a vector identifying a person using data characteristic of the face of this person. This database is stored for example in the working memory of the device 100.
[0066] At 402, the device 100 obtains a current image at time 't', denoted I_t as above, during a video call.
[0067] At 403, the device 100 implements a method for detecting people in the current image It. This detection provides as output, for each person detected, in particular one or more parameters making it possible to characterize each person and distinguish them from each other. The detection of people also provides information locating each person in the image.
[0068] According to one or more exemplary embodiments given without limitation, this detection comprises: - Body and face detection in the image - A face-body association - A vectorization of faces, the vectors being able to allow a comparison with the vectors in the database and thus classify them.
[0069] At 404, the face vectors are classified in the database containing the vectors of authorized persons. The classification indicates whether a person is authorized or not, that is to say whether they must remain present in the image or on the contrary be masked.
[0070] In 405, the device 100 constructs a mask for the image I_t, noted M_t, making it possible, once combined with the image It, to mask the people who must be masked, if such people have been detected.
[0071] According to one or more exemplary embodiments, the mask M_t is a function of: - Classification results in 404 - Output data from body and face detections in 403 - From the captured image I_t
[0072] At 406, the image S_t is generated as a function of the mask Mt, the image It and the image S_t-1, namely the modified image resulting from the previous iteration of the method of [Fig.4]. The image S_t is also recorded.
[0073] We then process the next image, returning to 402.
[0074] The different stages of the process of [Fig.4] will now be detailed.
[0075] Step 401 includes initializing a B_Temp database configured to allow classification of a person as an authorized or unauthorized person.
[0076] According to an exemplary embodiment, the database comprises for example a list of one or more authorized persons, with for each authorized person, a identifier 'i' and a vector 'D_ij' associated with person i, the vector containing the characteristics allowing person i to be distinguished from other people. A person not present in the list will by default be considered unauthorized.
[0077] According to an alternative embodiment, a person can be represented by several vectors (corresponding for example to different photos of the person's face).
[0078] According to an exemplary embodiment, the database B_Temp is initialized by copying the data from another database, B_Globale, present in the non-volatile memory of the device 100. B_Globale comprises, for example, a list of permanently authorized persons, as well as for each of these persons, a vector as described later.
[0079] The properties of vectors and their use for classification will be described in more detail later.
[0080] Step 402 comprises capturing a scene filmed by the camera 104. This capture is performed for example during a video call. An image is part of a sequence of video frames, each image in the sequence being processed successively within the framework of the method described, before transmission of the video consisting of the processed and, where appropriate, modified images to one or more recipients.
[0081] Step 403 concerns the detection of people in the image. It comprises the extraction of information necessary for the following steps, namely the information necessary for classifying the people present, and information helping, where appropriate, to mask, in the image, the people to be masked. According to an exemplary embodiment, this latter information describes the parts of the image occupied by a person.
[0082] Three sub-steps 501 to 503 will be described, namely the detection of bodies and faces (501), the body-face association (502) and the extraction of the parameters characterizing each person (503). These sub-steps are illustrated in [Fig.5].
[0083] Step 501 aims to detect the bodies and faces present in the image.
[0084] [Fig.6] is a schematic diagram illustrating the detection of bodies and faces according to a particular exemplary embodiment. [Fig.6] comprises an image It to be processed. Body detection delimits one or more zones C_i comprising bodies and face detection delimits one or more zones V_i comprising faces. In the example of [Fig.6], the zones are rectangular zones, also called bounding boxes, with V_i = {(X_vil, Y_vil),(X_vi2,Y_vi2)} and C_i = {(X_cil, Y_cil),(X_ci2,Y_ci2)} in the image frame.
[0085] According to a particular embodiment, algorithms known per se are implemented for this detection of bodies and faces. For the detection of both bodies and faces, the Viola-Jones method, also called 'Haar cascade' can be used. For face detection, the 'BlazeFace' neural network [1] can also be used. For body detection, the 'EfficientDet' neural network [2] can be used. Both neural networks output the bounding boxes C_i and V_i. The 'BlazeFace' network outputs the coordinates of the face bounding boxes, as well as the position of twelve landmarks (two for the mouth, four for the ears, four for the eyes, and two for the nose). The 'EfficientDet' network outputs the number and type of detected objects, and the bounding boxes in the image.
[0086] In the given example, an area delimiting a body contains the complete body, i.e. also the face.
[0087] Step 502 aims to produce a coherent association of the body 'C_i' and the face 'Vj' for the same person 'P_n'. We thus obtain sets P_n = {V_i, CJ}.
[0088] This association can for example be carried out on the basis of the zones detected during the previous step. According to a particular embodiment, the association of a face V_i and a body Cj comprises the calculation of the ratio between the intersection surface of the face V_i with the body Cj relative to the surface of the face V_i. A body Cj is associated with the face V_i with which it has the greatest ratio.
[0089] [Fig.7] is a schematic diagram illustrating the principle of body-face association.
[0090] According to an alternative embodiment, an additional condition is that the ratio is greater than a threshold. As an illustrative and non-limiting example, in certain applications, this threshold may be equal to 0.7.
[0091] As an example, the following pseudocode can be used to determine the association of a body with a face. Bodies are indexed with the index indx_c. Faces are indexed with the index indx_v, max_ratio represents the maximum ratio and max_indx represents the index of the face corresponding to the maximum ratio. max_indx and max_ratio are updated as the surface ratio is calculated for a given body during a loop in which each face is considered in turn. This loop over the faces is performed for each body.
[0092] 1 threshold = 0.7
[0093] 2 For indx_c, c in enumerate(body):
[0094] 3 max_ratio = -1, max_indx = -1
[0095] 4 For indx_v, v in enumerate(faces):
[0096] 5 ratio = surface (intersection (v,c)) / surface (v)
[0097] 6 if ratio > max_ratio && ratio > threshold:
[0098] 7 max_ratio = ratio;
[0099] 8 max_indx = indx_v;
[0100] 9 end if
[0101] 10 end for
[0102] 11 end for
[0103] At the end of step 502, each person P_n is then represented by a maximum of two bounding boxes, one for the face, the other for the body.
[0104] According to an alternative embodiment, in order to avoid associating the same face with several bodies, an associated face is excluded from the iterations for the following body(ies), i.e. once associated with a body, a face cannot be associated with another body.
[0105] According to an alternative embodiment, the case is considered where one or more faces are not associated with a body. This can happen for example if the detection of bodies and faces gives more faces than bodies. In this case, a person is represented only by a face, i.e. P_n = {V_i}. It is proposed to associate a fictitious body with such a face, so that the person concerned is represented both by a face and by a body, for the remainder of the method.
[0106] [Fig.8] is a schematic diagram illustrating the obtaining of a fictitious body for an unassociated face. In the example of [Fig.8], the dimensions of the fictitious body are a function of the size of the face. This is particularly simple to achieve in the case where the face and the body are considered to be enclosed in rectangular boxes. The dimensions of the bounding box of the face are noted w for the width and h for the height. The dimensions of the bounding box of the fictitious body result from the multiplication of the width w by a coefficient Kl and of height h by a coefficient K2. As a non-limiting illustrative example, Kl can be taken equal to 2 and K2 equal to 8. For example, if we denote by C' the fictitious body and (X Cji,Y'cji), (X'Cj2,Y'cj2) respectively the coordinates of the upper left point and the lower right point of the corresponding bounding box, we can for example obtain these coordinates by applying.
[0107] [Math.l] w = X^-X^
[0108] [Math.2] h =
[0109] [Math.3] y' _ y
[0110] [Math.4] Y'en = Yvil [YES] [Math.5] y' v JWnMMi) A c j2 ~ Xvi2 H--2-------
[0112] [Math.6] ^,2 = Y,n + K2x(Y„aY,,i
[0113] Other ways of determining a fictitious body can be considered.
[0114] According to an alternative embodiment, the surfaces corresponding to the bodies are detected, then for each body the corresponding face is detected inside the surface of the body. The face detected in the surface of the body is then directly associated with the corresponding body.
[0115] Step 503 comprises extracting a face from the parameters characterizing a person.
[0116] According to one embodiment, this step uses the principle of embedding, which comprises the generation of a vector of size N from an image in order to uniquely identify it. A distance calculation between two vectors, therefore two images, makes it possible to decide whether these two images are similar or not. According to a non-limiting embodiment, a cosine distance calculation is used. Other ways of calculating a distance between vectors can however be used. Two faces whose vectors are close in distance identify the same person.
[0117] A vectorization is performed for the faces V_i. This vectorization can be performed using tools known as such. For example, the facial recognition neural network present in 'Dlib' [3], a state-of-the-art library containing machine learning tools, can be used for vectorization. One implementation transforms a 150x150 pixel image into a vector of size 128.
[0118] At the output of this vectorization phase, each person P_i present in the scene is presented by two bounding boxes C_i and V_i, and a vector E_i of size N derived from the face V_i, as presented in [Fig.9], which is a schematic diagram illustrating the obtaining of two vectors El and E2 from respective faces VI and V2.
[0119] Optionally, a face bounding box is subjected to preprocessing before vectorization. This preprocessing consists of straightening or aligning the face using the landmarks in the face. This alignment makes it possible to obtain vectors at a shorter distance for different images of the face of the same person. The alignment consists of performing a transformation of the face image, for example a rotation, so that it is substantially vertical.
[0120] Returning to the method whose algorithm is given in [Fig.4], a classification of the faces is carried out in 404, including the determination, for each person detected, whether it is an authorized or unauthorized person.
[0121] According to one embodiment, this step comprises the use of the previously obtained Ei vectors. [Fig. 10] is a schematic diagram illustrating the principle of classification according to one or more embodiments.
[0122] The B_Temp database can be created in different ways and evolve over time. It should be noted that the different possibilities below are not mutually exclusive and can be combined in the same implementation. - The B_Temp database can be initialized with data present in another database (see above). - The B_Temp database can be expanded with people present at the beginning of the video call, for example, in the first thirty seconds. These people are then automatically authorized persons. For each of these people, one or more E vectors are stored in B_Temp.
[0123] These people are then authorized for the duration of the video communication, or according to an alternative embodiment, as long as they do not leave the filmed scene.
[0124] In the case where B_Temp is not initially empty, the device 100 determines for each vector E_i whether the database contains a vector close enough to conclude that the vector E_i corresponds to a person listed in the database. According to the present exemplary embodiment, to do this, the device 100 calculates the distance between each vector E_i and the vectors D_j already present in B_Temp. If for a vector E_i, a vector D_j is close enough - for example the cosine distance is less than a threshold e, (with for example e = 0.1) - the person "i" is considered authorized. Conversely, if for a vector E_i, no close vector is found in the database, then the person "i" is considered unauthorized.
[0125] Optionally, if a person is determined to be unauthorized, the user is asked if they wish to add this person as an authorized person in the B_Temp database.
[0126] Optionally, at the end of a video call, the user is asked if he or she wishes to add to the permanent database B_Globale one or more people present in the temporary database B_Temp but not yet present in the permanent database B_Globale.
[0127] Optionally, a user interface is provided so that a user can edit the B_Globale database, the editing including in particular the possibility of removing authorized persons.
[0128] According to an alternative embodiment, persons for whom no face is detected in 403 will be automatically considered as unauthorized.
[0129] The fact that all persons present in the database B_Temp are considered as authorized is not limiting: alternatively it is possible to implement a mechanism for constructing a subset P'K of authorized persons from the set of persons Pi present in the database so as to authorize only a subset of persons. This construction can be do so on the basis of one or more criteria, such as the type of communication, with certain people being indicated in the database as being authorized for certain types of communication only.
[0130] Steps 405 and 406 include processing the image It to make unauthorized persons invisible.
[0131] According to an exemplary embodiment, this processing comprises the creation of a mask (405) and the application of the mask to the image (406). Other implementations can be envisaged, in particular in a single step.
[0132] [Fig. 11] is a schematic diagram comprising an image (a), which is a current image I_t, an image (b) which is a mask resulting from a segmentation of the image (a) according to a first embodiment variant and an image (c) resulting from a second embodiment variant. The first embodiment variant comprises a semantic segmentation of the image and the second does not comprise any semantic segmentation.
[0133] A mask is a binary image, used to define a set of pixels of interest in an original image. The mask is for example defined by:
[0134] Mask (i,j) = 1 if the pixel Image(i,j) is a pixel of interest, and
[0135] Mask (i,j) = 0 otherwise.
[0136] In the present exemplary embodiment, the original image is the image I_t and the pixels of interest are the pixels corresponding to unauthorized persons. The mask has the same dimensions as the image I_t, but in other implementations, this is not necessarily the case. For example, the original image may result from a resizing of the image I_t, and the mask will then be smaller or larger in terms of pixels than the image I_t.
[0137] Mask construction step 404 implements the results of detection step 403 and classification step 404.
[0138] In the case of the first embodiment variant, a segmentation algorithm known per se can be used to construct the mask. The pixels of interest then relate fairly precisely to the part of the image corresponding to the person. This algorithm can be based on neural networks, such as for example the DeepLabV3 algorithm [4],
[0139] The second embodiment variant does not implement semantic segmentation. For example, the mask is then obtained by considering as pixels of interest the pixels of the bounding boxes corresponding to a person. This variant has the advantage of being less demanding in terms of computing resources.
[0140] The construction of a mask according to the first embodiment variant will now be described. The semantic segmentation comprises the association of a label with each pixel of the image. In the context of the present exemplary embodiment, the label of interest is the 'Person' label. [Fig. 12] is a schematic diagram illustrating the different stages of the construction of a mask implementing semantic segmentation according to an example of implementation.
[0141] Firstly, an extraction 1201 of the image areas containing the unauthorized persons from the original image is carried out. These areas are each placed in an intermediate image F_it, in this case F_2t in the example of the figure. The extraction is carried out using the coordinates of the bounding boxes of the face and body of these persons. The process first finds the coordinates of the extraction bounding boxes called “G_i” defined by:
[0142] X_lGi = min(X_cil, X_vil)
[0143] Y_lGi = min(Y_cil, Y_vil)
[0144] X_2Gi = max(X_ci2, X_vi2)
[0145] Y_2Gi = max(Y_ci2, Y_vi2)
[0146] where V_i and C_i are the bounding boxes of the face and body of person P_i classified as unauthorized. The size of the intermediate image F_it is then defined. In this implementation, F_it is an image where all pixels have the same color (white, in the example of [Fig. 12]). This image has for example an aspect ratio of 1:1 and a height which is equal to K*max((Y_2Gi - Y_lGi), (X_2Gi - Xl_Gi)). K can for example be taken equal to 2, as an illustrative example. The process then places the bounding box "G_it" extracted from the image I_t, in the center of the image F_it. The images F_it then pass to the segmentation step 1202 to generate the masks FM_it. The extraction bounding boxes "G_i" are then used in 1203 to extract the pixels of interest from the FM_it, and place them in M_t. Reference 1204 indicates the placement of the parts of interest of the FM_it in M_t at the same position as in the image I_t.
[0147] In the construction of a mask according to the first embodiment variant, the mask is constructed without semantic segmentation. According to this variant, the bounding boxes G_i, and constructs the mask M_t are obtained, by simply considering that all the pixels inside these boxes correspond to unauthorized persons and are therefore pixels of interest.
[0148] Once the mask M_t is constructed, the final processed image S_t can be obtained. [Fig. 13] is a schematic diagram that illustrates the composition of the final image.
[0149] The input data for this step includes:
[0150] - the generated mask M_t;
[0151] - the image S_t-1 sent at time t-1;
[0152] - the image I_t captured at time t.
[0153] To construct the image S_t, the following formula is applied: S_t = M_t * S_t-1 + (1-M_t) * I_t.
[0154] This formula means that:
[0155] 1. The value of pixel (i,j) of image S_t is equal to pixel (i,j) in image I_t if M_t(i,j) = 0.
[0156] 2. The value of pixel (i,j) of image S_t is equal to pixel (i,j) in image S_t-1 if M_t(i,j) = 1.
[0157] The image S_t is stored in volatile memory for the next iteration.
[0158] [Fig. 13] is a schematic diagram illustrating the composition of the image S_t.
[0159] According to one embodiment, the images S_t-1 are initialized (at t=0) with a image of the scene filmed by the camera without people. According to another embodiment, the people present at the start of the communication are authorized automatically. According to yet another embodiment, the initial image S_0 is simply a black image.
[0160] Removing a person from the image requires prior detection. Poor detection, or non-detection, can produce undesirable visual effects. One case where this problem may occur is when a person is partially visible in the filmed scene, for example when they are positioned on the edge of the image I_t and are only partially captured by the camera. [Fig. 14] is a schematic diagram illustrating the captured image I_t and the resulting image S_t, where a potentially unauthorized person straddles the edge of the image I_t and is not erased in the image S_t.
[0161] According to one embodiment, the image processing comprises a cropping which eliminates bands around the image to be transmitted, i.e. at least the lateral bands on both sides and in some embodiments also bands above and below. In the context of the examples previously illustrated, this cropping is therefore applied to the image S_t - this results in an image S'_t. The width of the removed bands is chosen so that a person entering the image does not appear in the cropped image - at least if the person enters at a certain distance from the camera. The width of the removed bands can also simply be a percentage of the dimensions, for example 5% of the width or height of the image. The cropping creates a larger detection margin, in order to increase the chances of good detection for people at the edge of the unprocessed image.
[0162] [Fig. 15] is a schematic diagram of a method for reducing the impact of the above-mentioned problem.
[0163] Images 1501 to 1503 illustrate an example in which it is not possible to detect a person and determine whether this person ('PA?') is authorized or not. In image 1501, this person is only located halfway in the image I_t. The processing described above is then applied to image 1501 to produce the processed image 1502. In the case of image 1501, the person at the edge of the image is not detected and is therefore not erased. Cropping is performed to eliminate at least the side bands. In the resulting image S'_t 1503, the undetected person does not appear. It is the image S'_t that will be transmitted.
[0164] Images 1504 to 1506 illustrate an example in which it is possible to detect the person entering the filmed scene. Image 1504 may correspond to the situation of image 1501 - the person has moved towards the center of the room. The detected facial area is then sufficient to determine the authorized / unauthorized status of the person. In the example of images 1504 to 1506, this person is not authorized. In the processed image 1505, the person will have been made invisible by applying the processing described above. However, the image is cropped to obtain an image S'_t 1506 in the same format as image 1503. If the unauthorized person enters the room even further, they will remain invisible in subsequent images S'.
[0165] According to a particular embodiment, the real background of the image, as filmed by the camera, is replaced by a virtual background. The processing applied is similar to that of [Fig.4], with the following modifications:
[0166] In step 404: - The construction of the mask is carried out with semantic segmentation. - The authorized person(s) become the elements of interest, and the masks are obtained for these people and not unauthorized people.
[0167] In step 405: - The composition of the image S_t is carried out using the mask M_t, the image captured at time t, I_t, and a virtual background image.
[0168] In step 406: - The transmitted image S_t does not need to be kept in memory for processing the next captured image.
[0169] According to an alternative embodiment, the possibility is given to switch between the real background of the camera image and a virtual background.
[0170] Facial landmarks are specific points on a human's face. These points are often placed around the face, eyes, and mouth. Locating such points can be accomplished using image processing methods known per se. The number of points used depends on the application and context. There are models, such as a model used by the previously mentioned 'Blazeface' algorithm, based on six points. The previously mentioned 'DLib' tool contains tools capable of using sixty-eight points.
[0171] As mentioned previously, in order to improve face vectorization, face alignment may be performed prior to vectorization. [Fig. 16] is a schematic diagram illustrating the alignment of a face according to one or more examples of implementation. The face in [Fig. 16] has twenty reference points. The principle of alignment consists of straightening the face to place on a line reference points symmetrical with respect to the vertical axis of symmetry of the face, for example points 2 and 6 above the eyes, or reference points that must be on the axis of symmetry of the face, such as points 17 and 20. From such points, an angle of rotation alpha of the face with respect to the vertical is determined. The image of the face is then transformed by an operation of rotation of angle alpha to straighten it vertically. List of documents cited
[0172] [1] Bazarevsky et al., 'BlazeFace' - Article “BlazeFace: Sub-millisecond Neural Face Detection on Mobile GPUs“ ('BlazeFace: Sub-millisecond Neural Face Detection on Mobile Graphics Processing Units') - July 14, 2019 -http: / / arxi.org / abs / 1907.05047
[0173] An implementation is available in https: / / developers.google.com / mediapipe / solutions / vision / face_detector
[0174] [2] Tan et al., 'EfficientDet' - Article “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks” ('EfficientDet: Rethinking Model Extension for Convolutional Neural Networks') - Sept. 11, 2020 - https: / / arxiv.org / abs / 1905.11946
[0175] [3] 'DLib'https: / / github.com / davisking / dlib-models, and in particular dlib_face_recognition_resnet_model_vl for embedding and shape_predictor_68_face_landmarks for facial landmark detection.
[0176] [4] Chen et al., 'Rethinking Atrous Convolution for Semantic Image Segmentation' ('Rethinking atrous convolution for semantic image segmentation') - Dec. 5, 2017 - https: / / arxiv.org / abs / 1706.05587
[0177] An implementation of DeepLabV3 is available in https: / / tfhub.dev / tensorflow / lite-model / deeplabv3 / l / metadata / 2
Claims
Claims
1. A method of video communication implemented by a device (100) comprising a processor (107) and a memory (105) comprising software code, the processor executing the software code causing the device to implement the method, the method comprising: - obtaining (201) an image (I_t) generated by a camera; - detecting (202) one or more persons (PA, PN) in the image; - if one or more persons are detected, checking (203) for each person a criterion indicating whether the person should or should not be part of the video communication; - processing (204) the image to erase from the image the detected person(s) not to be part of the video communication; - transmitting (205) the processed image (S_t, S'_t).
2. Method according to claim 1, wherein the detection comprises the identification (501) of one or more first zones of the image (V_i), each first zone comprising a face.
3. A method according to claim 2, wherein the detection further comprises: - identifying (501) one or more second areas of the image (C_i), each second area comprising a body; - associating (502), based on an association criterion, a first area and a second area to form a representation of a person (P_i) in the image.
4. A method according to claim 3, comprising, for a given first area not being able to be associated with a second area on the basis of the association criterion, determining a third area (C'_i), where the third area is an area of the image which is a function of the given first area and is intended to serve as a second area associated with the given first area to form a representation of a person in the image for image processing.
5. A method according to either of claims 3 or 4, comprising, when a
6.
7.
8.
9. second zone cannot be associated with a first zone, the marking of this second zone as part of a person who must not participate in the video communication, the representation of this person then comprising only the second zone. Method according to one of claims 3 to 5, comprising the extraction (503), from each first zone, of characteristic parameters (E_i) of the face of each first zone, said characteristic data being adapted to make it possible to determine, from a database, whether a person corresponding to a face must or must not be part of the video communication. Method according to claim 6, said database comprising: - either characteristic face parameters for one or more persons authorized to participate in the video communication; or at least one of: - characteristic parameters for one or more faces and for each face an indication that the person corresponding to the face should be part of the video communication; and - characteristic parameters for one or more faces and for each face an indication that the person corresponding to the face should not be part of the video communication. The method of claim 7, comprising initializing (401) the database with pre-stored data. A method according to claim 7 or 8, comprising augmenting the database by: - the identification of one or more first zones comprising a face, during a time interval from the start of the video communication; - the extraction, from each first zone, of characteristic parameters (E_i) of the face of each first zone; - the storage of the characteristic parameters of the face of each first zone with a respective indication that the person matching the face must be part of the video communication.
10. Method according to one of claims 3 to 9, the processing of the image comprising - obtaining a mask (M_t) from the representations of persons not to be part of the video communication; - obtaining an image processed according to the mask, the image (I_t) obtained from the camera and a processed image (S_t-1) obtained previously.
11. Method according to one of claims 1 to 10, comprising, following the processing of the image to erase from the image the detected person or persons not to be part of the video communication, a cropping of the processed image (S_t) to obtain a cropped processed image (S'_t), the cropping being configured to remove at least left and right side bands from the processed image, the side bands being of a width capable of containing one or more persons before their detection becomes possible, the transmission being carried out with the cropped processed image.
12. A device (100) comprising a processor (107) and a memory (105) comprising software code, the processor executing the software code causing the device to implement the method according to one of claims 1 to 11.
13. A television decoder comprising a device according to claim 12.
14. Computer program product comprising instructions which when executed by at least one processor cause the execution of the method according to one of claims 1 to 11 by said at least one processor.
15. A non-transitory storage medium comprising instructions which when executed by at least one processor cause the method according to one of claims 1 to 11 to be executed by said at least one processor.
Citation Information
Patent Citations
Method and apparatus for isolating an active participant in a group of participants
EP3101838A1
Cinematic image framing for wide field of view (FOV) cameras
EP4080446A1
Context based target framing in a teleconferencing environment
US10904485B1
Automatically enhancing privacy in live video streaming
US20190147175A1
Secure nonscheduled video visitation system
US20230179741A1