Image data processing, particularly from video conferencing, to determine the interlocutor of a speaking speaker

The method determines the local interlocutor in videoconferencing by analyzing head orientations, addressing the challenge of identifying the local speaker in complex multi-participant settings, enhancing videoconferencing efficiency and enabling accurate transcripts.

FR3163196A1Pending Publication Date: 2025-12-12ORANGE SA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
FR2024006190
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-11
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Current videoconferencing technologies struggle to identify the local interlocutor of the current speaker when multiple participants are present at the same site, leading to frustration and reduced effectiveness in multi-site meetings.

Method used

A method for determining the local interlocutor of the current speaker based on the head orientations of multiple users, using image data processing to estimate angular deviations and assign weights to the current speaker, with the option of facial recognition for identity confirmation.

Benefits of technology

Enhances videoconferencing efficiency by providing information on both the current speaker and the local interlocutor, allowing for improved video transmission and verbatim transcripts, and can be applied to other scenarios like filmed debates or multimedia audio description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method is proposed for processing image data of a scene with multiple users, including a current speaker at a given moment. The method involves determining, among these users (excluding the current speaker), an interlocutor of the current speaker at that given moment, based on at least the respective head orientations of at least some of the users. (See abstract figure: Figure 4)
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Processing of image data, particularly video conferencing data, to determine the interlocutor of a speaking speaker technical field

[0001] This disclosure falls within the field of image data processing, particularly videoconferencing. Previous technique

[0002] Videoconferences held between a plurality of sites often require equipment including a mobile (motorized) camera and a screen, at each site.

[0003] The screen displays the videos captured by the cameras of the remote site(s): each user can thus view a mosaic of the remote sites.

[0004] A multi-site meeting becomes more efficient if the display of the current speaker is well managed (typically with their video highlighted in a brightly colored box, for example), to encourage listening to the current speaker by the other sites.

[0005] When a site hosts a single user, that user typically looks at the camera naturally when addressing other remote participants. Conversely, the situation is more complex when several participants occupy the same site: - the speaker may address the remote participants, and the speaker's gaze is then turned towards the camera, or - The speaker is addressing a specific person present at their location and their gaze is directed towards that person.

[0006] This second situation involving a local interlocutor is much more complex for remote participants to grasp. This is precisely a difficulty that current videoconferencing solutions cannot overcome, leading to both frustration for remote participants and reduced videoconferencing effectiveness.

[0007] In a nominal situation, both the current speaker and the local interlocutor can be in the camera's field of view, and remote participants then normally have the opportunity to guess who the local interlocutor is by analyzing the current speaker's gaze direction. However, in more complex situations (for example, if the camera zooms in on the image of the current speaker), it is no longer possible to guess the local interlocutor. Summary

[0008] This disclosure improves the situation.

[0009] A method for processing image data of a scene with multiple users, including a current speaker at a given moment, is proposed. The method specifically includes determining, among said users other than the current speaker, an interlocutor of the current speaker at said given moment, based at least on the respective head orientations of at least some of the users.

[0010] Thus, the treatment proposed here is not based solely on the orientation of the current speaker's head, who may look elsewhere than towards his interlocutor (for example towards the ceiling or towards the screen, or something else), but on a larger number of users participating in the aforementioned scene.

[0011] Such processing to determine the local interlocutor of the current speaker is particularly useful for the efficiency of a videoconference, for example. The video transmission can then be enhanced with information concerning not only the current speaker but also the local interlocutor. Furthermore, for a text transcript of the meeting, this information can be accessible (for example, in verbatim transcripts such as: "Mr. X addressing Ms. Y: xxx").

[0012] In one embodiment, the image data can thus be videoconference data. However, other applications of the process besides videoconferencing can be envisaged, in particular for transcribing the verbatim transcripts of a debate or any scene that has been filmed.

[0013] In one embodiment, the aforementioned determination includes an estimation of angular deviations between: - a head orientation of each user of said user group, and - a direction defined by a position of that user towards a position of the interlocutor.

[0014] This interlocutor may be a presumed interlocutor if it has been detected in one or more previous images and the aim here is to confirm it as an interlocutor. In this case, for example, if the estimate of the angular deviations (for example, the sum or average of the angular deviations) is less than a threshold, then the interlocutor can be confirmed.

[0015] Alternatively, it may be a candidate user as an interlocutor of the current speaker, among several candidates.

[0016] Indeed, in an embodiment, for each candidate as an interlocutor of the current speaker among the users other than the current speaker, the method may include an estimation of an average of the angular deviations between: - a head orientation of each user of said part of the users, and - a direction defined by a position of this user towards a position of the candidate, at least one candidate being retained as a possible interlocutor of the current speaker if the estimate of the average angular deviations for that candidate is less than a threshold.

[0017] In such an embodiment, a weight may be assigned to the current speaker that is greater than the weight of another user in said part of the users, in the estimation of said average.

[0018] For example, this weight can be three or four times that of another user than the current speaker. Of course, this value is configurable, for example according to the number of participants in the scene, or other factors.

[0019] In an embodiment, the candidate whose average estimate is minimal among the average estimates of the candidates can be retained as a possible interlocutor of the current speaker.

[0020] In one embodiment, particularly for determining the respective positions of users, the method may include: - detect the respective faces of users in the scene, - assign unique identifiers to each user whose faces are detected, and - determine current positions of each user in the scene corresponding to said identifiers, in a top view coordinate system of the scene.

[0021] Face detection here is not necessarily facial recognition of specific faces. It is sufficient to detect objects as human heads and to assign them, on the one hand, an identifier (any index for example), and, on the other hand, positions associated with these identifiers.

[0022] Such an implementation allows, for example, the various average calculations mentioned above to be carried out, by considering in turn the users present in the scene as possible candidates.

[0023] In one embodiment, however, a provided identity of the users may be taken into account, particularly for other purposes, and the method may further include: - providing image data of the faces of users present in the scene, corresponding to the respective identifiers provided for said users, - implementing recognition of the respective faces of the users in the scene, and - determine the current positions of users in the scene in correspondence with the respective user IDs provided.

[0024] In one embodiment, the process may further comprise: - to memorize the identifier of the selected candidate as a possible interlocutor for the current speaker, corresponding to a given moment, - estimate a score, from a weighted average with a forgetting factor, over a sliding time window including the given instant and one or more instants preceding the given instant, assigned to the memorized identifiers of the selected candidates, for each selected candidate at the given instant and at the preceding instants, and - determine the selected candidate with the best score as an interlocutor of the current speaker for the given instant.

[0025] Such an achievement then makes it possible to "smooth over time" the determination of the interlocutor according to past and present determinations.

[0026] In addition, in the calculation of this "temporal" average, a weight greater than the weights of other selected candidates may be assigned to a selected candidate who was a fluent speaker at at least one of the said preceding moments.

[0027] Indeed, if it is typically a natural dialogue, it is likely that the speaker for past moments will become an interlocutor of a current speaker.

[0028] In an embodiment where the scene is a videoconference room with a screen, it is determined that the possible interlocutor of the current speaker is in a remote site participating in the videoconference if the respective head orientations of the users in the room are detected as directed towards the screen, for a number of users greater than a threshold (for example 60% of the users, the current speaker being able to have more weight).

[0029] Alternatively, the screen can be considered a candidate in the same way as a user and its weighted average is estimated as described above: the screen can thus be retained if its average is minimal among all the candidates.

[0030] In addition, if it is determined alternatively that the possible interlocutor of the current speaker is in the videoconference room, users whose head orientations are detected as directed towards the screen may be excluded from said part of the users (because they may unnecessarily distort the determination by the aforementioned calculations of angular deviations).

[0031] In addition, users whose head orientations are detected as directed (naturally) towards the current speaker can be excluded from said part of the users, because of course, the current speaker cannot be his own interlocutor.

[0032] In a particular embodiment, the process may include: - provide image data of users' faces present in the scene, corresponding to the respective identifiers provided for said users, - Implement facial recognition for each user in the scene, - Store the identifier of the user identified as the current speaker's interlocutor at a given moment, with a view to transmitting said stored identifier, corresponding to that moment, to at least one script editor for display said identifier memorized in verbatim transcripts of everyday speakers and interlocutors at respective given times.

[0033] In addition or alternatively, the scene is a videoconferencing room comprising equipment connected to at least one remote videoconferencing device, and the method may include: - detect the respective faces of users in the scene, and - to store at least some data from a position of the user determined as the interlocutor of the current speaker at a given time, with a view to transmitting the data of said position, corresponding to the given time, to the remote video conferencing equipment, configured to drive a display of the image of the scene at the given time with highlighting of the interlocutor (for example by producing an additional image with a digital zoom on the position of the interlocutor, or by embedding in the video conferencing image a bright coloured frame around this position).

[0034] It should be noted that in such an embodiment, the exact identity of the interlocutor is not necessary and does not need to be transmitted. Moreover, facial recognition of the participants, and in particular of the interlocutor, is also unnecessary.

[0035] According to another aspect, a computer program is proposed comprising instructions for implementing all or part of a process as defined herein when this program is executed by a processor. According to another aspect, a non-transient, computer-readable recording medium is proposed on which such a program is recorded.

[0036] According to another aspect, a device for processing image data of a scene comprising a plurality of users including a current speaker at a given time is proposed, comprising a processing circuit for implementing the above process. Brief description of the drawings

[0037] Other features, details and advantages will become apparent from reading the detailed description below and from analyzing the accompanying drawings, in which: Fig. 1

[0038] [Fig.1] shows an example of a succession of general steps of a process of the type described above. Fig. 2

[0039] [Fig.2] illustrates a possible alpha angle estimation by image analysis between a viewpoint and a particular object OBJ of a scene filmed by a camera having a shooting angle FOV. Fig. 3

[0040] [Fig.3] illustrates how algorithms trained to predict relative proximity can provide a depth estimate (pixel by pixel) without even detecting or locating objects or people. Fig. 4

[0041] [Fig.4] illustrates a videoconference room seen from above and comprising users participating in the videoconference, as well as videoconferencing equipment comprising a CAM camera and an ECR screen, and connected to a DIS device for the implementation of the method, according to one embodiment. Fig. 5

[0042] [Fig.5] illustrates a landmark historically and traditionally used to describe the orientation of the head. Fig. 6

[0043] [Fig.6] shows a succession of steps of a process of the type presented above, according to one embodiment. Description of the implementation methods

[0044] It is proposed to determine the local interlocutor (hereinafter referred to as the "target interlocutor" or "interlocutor") of a current speaker in a videoconference session, using a camera (for example, a simple monocular camera). With reference to [Fig. 1], a method for such determination may, for example, comprise seven main steps.

[0045] During a first PI step, the process can be triggered by manual intervention, or automatically by a request from another site, on a regular basis, for example. In one embodiment, this PI step of triggering the process is implemented automatically and repeatedly without human intervention.

[0046] The second step, P2, aims to detect the faces of different people present on site. Neural network algorithms can perform multi-object detection in an image, including human shapes or faces. This detection can be robust, particularly in situations of partial occlusion. The topology of a meeting room often facilitates detection because the furniture is precisely arranged to limit occlusions. This does not necessarily involve recognizing specific faces, but simply detecting objects such as human faces.

[0047] The third step P3 aims to estimate the different positions of the detected faces. Even though, in the previous step P2, a neural network was able to determine the presence of different people on the site, such detection, which can be defined as a bounding rectangle in the overall image, does not allow for...

[0048] To deduce the location of an object at an absolute coordinate within the scene, two techniques can be used together: - a first technique consisting of determining an angle between the detected object and the camera axis ([Fig.2] discussed later), and - a second technique consisting of determining the distance between the object and the camera ([Fig.3] discussed later), which typically allows the transition from a cylindrical coordinate system to an orthonormal system. Figure 2 illustrates an example of parameters for determining the aforementioned angle (denoted alpha in the figure). As a reminder, the camera is positioned in front view. It is necessary to know the maximum angle of the camera's field of view, denoted FOV (for Field of View). The unknown is the angle alpha. In pixel space, the total width of the image, denoted w below, is known. Regarding the detected object OBJ, its center can be considered as illustrated in the figure, and the deviation in pixels from the central axis of the image is observed. This central deviation is denoted AI (delta-I) in Figure 2. The following relationship between these parameters can then be shown: —2— — AL FOVn w / 2

[0049] Distance estimation (the second technique mentioned above) can be implemented using artificial intelligence, simply from a frontal view. Indeed, approaches to estimating (relative) proximity can be robust using only a single monocular view. Thus, using a single image, such artificial intelligence can provide a table of values ​​corresponding to an inference result for each pixel. As an example in [Fig. 3], the table of relative proximities is presented on the right as a grayscale image (with brightness proportional to proximity), based on the actual frontal image on the left. Typically, the space under the tables of people in the front row appears brighter, because it is closer to the camera, than the background (in black), which is much farther away.

[0050] Such an approach has the advantage of using a standard, low-cost camera. It is therefore unnecessary to use a stereoscopic camera or one equipped with a sensor such as a Lidar. However, it is possible to avoid using standard cameras with optics that induce strong distortions, such as wide-angle lenses (of the "fisheye" type, for example).

[0051] It may be noted that the distance "capture" operates within a scale factor, but this is by no means a limitation for the implementation of the following method. Indeed, as suggested by the top view diagram in [Fig. 4], the approach is purely geometric and based on angles: it is precisely independent of a scale factor.

[0052] The next step, P4, aims to capture the orientation of faces. The established term in the English-language literature is "Head Pose Estimation." More generally, this step P4 consists of capturing the head orientation of all participants in the meeting room (including those who may have their backs to the camera).

[0053] Head orientation from a single image can be determined using, for example, a pre-learning head orientation prediction model. Most models of this type perform face and facial marker detection (eyes, nose, mouth, etc.) to facilitate learning. These approaches are then only valid for the interval [-90°, +90°]. Other approaches bypass facial markers to rely directly on head detection (without using the face) and thus provide an estimate for a wider interval. A difficulty then arises related to the discontinuity of the rotation angle from +180° to -180°, which machine learning models struggle to handle.

[0054] To address these problems, it is proposed to change the mode of representation. Euler angles have the dual advantage of conciseness and ease of interpretation. A general rotation matrix does not offer these advantages but provides a quality of continuity that is compatible with the effective use of machine learning algorithms. Considering, for example, orthonormal matrices, the last column is deduced from the first two, so that only six parameters remain to be determined. The use of this representation offers better results, and a transcription into Euler angles is possible (for example, in a top-view coordinate system as shown in [Fig. 4]).

[0055] It is then necessary to distinguish the speaker among the participants in the room. Prior art techniques make it possible to determine the speaker. Moreover, many videoconferencing services offer a motorized camera that focuses precisely on the speaker. The identification of the speaker in the image is therefore information produced by some process of this type of prior art and is not detailed here.

[0056] With a high probability, the current speaker's head is turned towards their interlocutor. As for the participants, their heads are primarily turned either towards the current speaker or towards their interlocutor. For the next step P5 (determining angular differences in pairs), a combinatorial set optimization is proposed that determines, among the n-1 participants, which one has the highest probability of being the aforementioned interlocutor. This is a two-dimensional geometric approach (top view conforming to [Fig. 4]) based on the location of each participant and the orientation of their head.

[0057] To this end: - an algorithmic loop checks the probability of these n-1 hypotheses, - in each loop, on the n-2 remaining participants (having removed the current speaker and the candidate): * We calculate the absolute angular difference between a "gaze" direction (head-first, for example) of this participant and a candidate participant acting as an interlocutor of the current speaker. * We also determine the angular difference, in absolute value, between the participant's gaze and the direction towards the current speaker, * we also determine the angular gap, in absolute value, between the participant's gaze and the screen (as another possible source of direction of a participant's gaze).

[0058] Calculating these last two values ​​allows us (with detection of small angles) to determine whether the participant is clearly looking at another participant (possibly the interlocutor), or rather (naturally) at the current speaker, or even at the screen. Only in the first case is the absolute angular difference included in the average calculated to support the hypothesis of the candidate interlocutor. The angular difference between the current speaker and the candidate interlocutor is also included, with a higher weighting coefficient, for example x3, or another.

[0059] Such an achievement amounts to carrying out a form of "vote" in which the votes corresponding to the participants who are potentially looking (within an angular margin of error) at the candidate interlocutor are counted, giving more importance to the vote of the current speaker.

[0060] The most probable hypothesis corresponds to a minimum of this weighted average. Thus, the interlocutor of the current speaker can be identified in the next step P6 as having the lowest average among all the candidates.

[0061] This implementation can be carried out using a single image. In one embodiment, it is possible to take advantage of the temporal continuity from one image to the next, due to very little movement of the people and minimal changes in head orientation (particularly in a meeting setting where participants are presumably seated and move relatively slowly). A standard camera typically captures thirty images per second. Even if, for the sake of simplicity, only a few images per second are processed, images very close together in time reinforce each other, especially for finding the target speaker: the intended speaker is more likely to be the same person in this closely spaced sequence. It is then proposed to perform a set-theoretic optimization combining the location of the participants, head orientation, and temporal sequences.Thus, steps P2 to P6 can be repeated (over a given time window), then the estimates of "votes" or scores of candidate interlocutors can be averaged at the end of the step. P6 (on a sliding window with a forgetting factor) in order to arrive at a definitive interlocutor identification on this window.

[0062] Furthermore, in an embodiment where several cameras are used to acquire images of the meeting, an average of the scores established from the data acquired by each camera can be applied in the same way. An additional operation of associating the faces detected from one camera view to another can be added, and the determination of the speaker is derived from the average (or a majority vote, or other method) of the speaker determinations from all viewing angles.

[0063] It should be noted that in the case of a meeting with remote participants, the presumed speaker may simply not be physically present in the room. In this case, the participants in a room, as well as the speaker, tend to look at the screen displaying the videoconference (and therefore towards the camera). The speaker is then considered to be remotely connected. In one embodiment, it is possible to determine which remote person (and / or in which specific remote room) is likely the speaker by carefully analyzing the direction of the participants' gaze according to the fixed position of that person / seat on the videoconference screen.

[0064] The next step P7 can then involve the transmission of data relating to the interlocutor of the current speaker (for example, their name or identifier in the videoconference session). The interlocutor's identification information can be transmitted to a third-party system.

[0065] For example, the visual interface displayed on a third-party videoconferencing screen can then be enhanced with a live image of the other participant generated by the videoconferencing software. For example, the video feed from the room where the participant is located can be split in two on the third-party site's screen, with a zoom on the current speaker on one side and a zoom on the other participant on the other. Typically, in this embodiment, facial recognition of the current participant is not required. Their position POS_i has been determined based on the identifier ID#i that was arbitrarily assigned to them (i.e., without knowing the participant's specific identity). Thus, it is possible to transmit a digital zoom (in the acquired image) of this POS_i position to the remote site(s) or to embed in the image a frame around this POS_i position highlighting the interlocutor.

[0066] Furthermore, it may be possible to display, during a meeting transcript, the identity of the person being addressed (if facial recognition has been implemented), for example, in the form "Mr. X addressing Mr. Y: xxxx". Thus, the transcript of a videoconference in the case of a given session may display the name and / or a photograph of the speaker and their Participant photographs may be available before the videoconference session, or taken during the session. Furthermore, previously available photographs can be used to facilitate the aforementioned facial recognition of non-speaking participants during the session.

[0067] Thus, the proposed image analysis techniques can make it possible to obtain data on the orientation of participants' gazes towards a local interlocutor of a fluent speaker. This data can be determined from a simple camera.

[0068] Indeed, a first general step is to capture the location of all meeting participants simply through the frontal view provided by a camera mounted on the wall. This can be a standard, ordinary, and low-cost camera without resorting to a stereoscopic (multi-lens) camera or one equipped with distance sensors. Artificial intelligence can be implemented to first detect the participants' faces, and then, in a second general step, the orientation of their heads and / or their gaze. As illustrated in [Fig. 4], typically, in a top-down view of a videoconferencing room, knowing the participants' locations and the direction of their gaze, it is possible, knowing the current speaker, to determine the local interlocutor to whom the current speaker is addressing themselves. The aforementioned second general step aims to capture the orientations of the faces to allow the calculation of the most probable target interlocutor.It is not necessarily necessary for the face to be fully visible to determine its orientation. An estimate of the head's orientation can be made, even from the back.

[0069] More precisely, an angular deviation can be measured between the orientation of the speaker's head (arrow from the speaker in [Fig. 4]), for example, and a straight line (dotted lines from the speaker) passing through the speaker's position, on the one hand, and that of a candidate about whom it is hypothesized that this candidate might be the speaker's interlocutor. It is thus understood that the smaller this angular deviation, the more well-founded the hypothesis that this candidate is the speaker's interlocutor.

[0070] However, as also illustrated in [Fig. 4], where the speaker does not direct their gaze directly at their interlocutor but rather to the side of the interlocutor, an approach focused solely on the speaker's gaze direction could, in some cases, present limitations and weaknesses. Therefore, in one particular embodiment, it is proposed to strengthen the approach with a comprehensive analysis of the gaze directions of the local participants ("Angle 1," "Angle 2," "Angle 3," in [Fig. 4], in addition to the speaker's angle), who generally tend to look at the speaker or their interlocutor. This comprehensive information works together to determine, with increased probability, the target interlocutor for the conversation.

[0071] Furthermore, a temporal analysis of the sequence of images helps to consolidate this determination, giving more weight to a previous speaker because the current speaker is more likely addressing the person who spoke immediately before them. The target interlocutor is, in fact, with a high probability, this previous speaker.

[0072] Figure 6 illustrates a particular embodiment, by way of example, for determining the current speaker's interlocutor. In a first step, SI, photographs of the people present in the videoconference room are obtained, corresponding to a name or identifier for each of these people. This embodiment then allows names to be assigned to the various meeting transcripts. Furthermore, information is obtained about the current speaker among these people.

[0073] In step S2, within a current IM image, the faces of the people present in the videoconference room are detected (including the face of the current speaker LC). From this, the respective positions POS_i of each of these people are deduced, with their respective identifiers ID#i. These faces can also be recognized from the aforementioned photographs (by facial recognition), and in this case, the identifiers ID#i assigned to these people can be more specific (names, surnames, nicknames, or other identifiers).

[0074] The coordinates of the POS_i positions, for example in a frame of reference which includes an image taken by the CAM camera, are transferred to a frame of reference in a top view as illustrated in [Fig.4], in the next step S3.

[0075] In addition, the angular orientations of the heads of the participants ORI_i are determined in this lair (see the arrows from the participants on [Fig.4]).

[0076] Next, in step S5 (in the illustrated embodiment), head orientations are sought that are turned towards the current speaker (LC) or towards the ECR screen of the videoconference room. Such an embodiment makes it possible to avoid involving people who are naturally looking at the current speaker (LC), or even the videoconference room screen, in determining the interlocutor of the current speaker.

[0077] Thus, in step S6, for each candidate m among the videoconference participants, it is determined whether the other participants, in their respective positions POS_i, have their gaze ORI_i directed towards this candidate m. Excluded from this determination are, of course, candidate m himself, as well as those who have been detected as looking at the screen (head orientation towards the screen ECR). If, for example, the majority of people detected as looking at the screen are considered to be the interlocutor of the current speaker, it may be decided that the interlocutor of the current speaker is ultimately a person from another site participating in the videoconference.

[0078] Otherwise, to carry out the determination of step S6, an angular deviation is calculated between the head orientation ORI_i of a selected participant, and a straight line passing through the respective positions of the selected participant POS_i and the candidate POS_m. More specifically, an average (or sum) of these angular deviations MOYm is calculated for each candidate m. As it is likely that the current speaker LC is speaking fluently to their interlocutor, a greater weight is assigned, in the calculation of this average, to the participant who corresponds to the current speaker LC.

[0079] In the next step S7, the candidate m whose mean MOYm is the smallest among the calculated means can be designated as the interlocutor of the current speaker INTLO(t) for an image IM acquired at a time t.

[0080] Of course, this determination of the current speaker's interlocutor can be "smoothed" over time. For example, at step S8, a forgetting factor can be determined for previous interlocutor determination times INTLO(t). For example, an average over time can be calculated, with a weighting of the type: - weight of the determination of INTLO(tn) = 1, - weight of the determination of INTLO(tnl) = 2, - weight of the determination of INTLO(tn-2) = 3, ... - weight of the determination of INTLO(t) = n.

[0081] Furthermore, for the calculation of this time average, in step S9, if among the previous times of determination, the current interlocutor INTLO(t) had been the current speaker LC(tk), then it assigned more weight to the present determination than the interlocutor INTLO(t). For example, weights of the type: - weight of the determination of INTLO(tn) = 1, and =K if INTLO(t) = LC(tn) - weight of the determination of INTLO(tnl) = 2, and =K (or K+l) if INTLO(t) = LC(tn-1) - weight of the determination of INTLO(tn-2) = 3, and =K (or K+2) if INTLO(t) = LC(tn-2), ... - weight of the determination of INTLO(t) = n.

[0082] It will thus be understood that the identifiers assigned to these interlocutors at previous times are stored in memory. It is then possible to re-identify a person from one image to another to establish the correspondences between past and present interlocutors. The person detection algorithms (human form or face) can then provide associated semantic information for each detected object, which notably allows for matching people between the different images, both previous and current.

[0083] Of course, for the choice of values ​​to be assigned to all the aforementioned weights, the process presented here is configurable, in particular according to the context of the ongoing videoconference (according to the number of participants in the same room, and / or the number of sites participating in the same videoconference, and / or the total number of participants, ...).

[0084] At the end of the processing, in step S10, an interlocutor of the current speaker INTLO(t) is thus determined for a current time t. As described previously with reference to [Fig. 1], data on the interlocutor's position POS_i in the acquired image can be stored in memory and then transmitted to the remote site to perform a digital zoom in the received image (for example, in a second image displayed on the remote site) or to highlight the interlocutor with a box around their position POS_i. If photographs of the participants are provided in step S1, their specific identifier (surname, first name, etc.) can be transmitted as a supplement or alternative for displaying their identifier in the videoconference image displayed by one or more other sites, or for tracking the verbatim transcripts of the videoconference meeting.

[0085] Figure 4 also illustrates a possible device for implementing the above processing, as defined in Figure 6, or more generally in Figure 1. Such a DIS device is connected (via radio frequency or wired connection) to videoconferencing equipment located in the room. Such equipment typically includes, in particular, the ECR screen, the CAM camera, a videoconferencing data encoder / decoder (not shown), and a connection to a wide area network.

[0086] The DIS device for implementing the processing may, for its part, comprise: - a MEM memory storing, in particular, instruction data from a computer program for implementing the aforementioned processing, and possibly photographic data of the participants PIC#i corresponding to their name or ID#i identifier as described above with reference to step SI of [Fig. 6]; - a PROC processor capable of cooperating with the MEM memory to read the aforementioned instruction data and execute processing in accordance with the computer program; and - a communication interface between the PROC processor and the video conferencing equipment.

[0087] Of course, this DIS device can be in the form of a module integrated directly into a processing circuit of the aforementioned videoconferencing equipment.

[0088] Among the application areas of such a development, the instrumentation of hybrid videoconferences is particularly noteworthy. Current services correctly capture the current speaker. The processing proposed above captures the interlocutor of this current speaker, providing a differentiating advantage for videoconferencing services, which are constantly gaining importance in the evolving world of work.

[0089] The proposed processing can be applied to image data acquired by any type of camera, including webcams and those of other mobile electronic devices equipped with at least one photo sensor, such as a standard mobile phone, provided that the optics are arranged so as to have an overview of a meeting room for example.

[0090] Of course, the present description is not limited to the embodiments described above by way of example; it extends to other variants.

[0091] As described previously, a transcript of the verbatim transcripts of a videoconference meeting can be implemented using the processing method proposed above. However, other use cases besides videoconference meetings can be considered, such as the transcription of interactions within a group of people, for example, a filmed debate, a play, or a film. As presented above, a directory of participants with their photographs allows for the capture of the speakers' identities, and from there, automated transcription of the dialogues ("X is speaking to Y"), particularly for generating multimedia audio description data for films or videos more generally.

Claims

Demands

1. A method for processing image data of a scene comprising a plurality of users including a current speaker at a given time, the method comprising determining, among said users other than the current speaker, an interlocutor of the current speaker at said given time, based on at least orientations (P4; S4) of the respective heads of at least some of the users.

2. Method according to claim 1, wherein said determination includes an estimation of angular deviations (P5; S6) between: - a head orientation of each user of said user part, and - a direction defined by a position of this user towards a position of the interlocutor.

3. A method according to claim 2, comprising, for each candidate as an interlocutor of the current speaker among the users other than the current speaker, an estimate of an average (S6) of the angular deviations between: - a head orientation (ORI_i) of each user of said part of the users, and - a direction defined by a position of this user (POS_i) towards a position of the candidate (POS_m), at least one candidate being retained as a possible interlocutor of the current speaker if the estimate of the average of the angular deviations for this candidate is less than a threshold.

4. Method according to claim 3, wherein a weight is assigned to the current speaker higher (S6) than a weight of another user of said part of the users, in the estimation of said average.

5. A method according to any one of claims 3 and 4, wherein the candidate whose mean estimate is minimal (S7) among the mean estimates of the candidates is retained as a possible interlocutor of the current speaker.

6. A method according to any one of the preceding claims, comprising: - detecting (P2; S2) the respective faces of the users in the scene, - assigning unique identifiers (ID#i) to each of the users whose faces are detected, and - determine (P3 ; S3) current positions (POS_i) of each user in the scene corresponding to said identifiers, in a top view coordinate system of the scene (S3).

7. A method according to claim 6, further comprising: - providing (S1) image data of user faces present in the scene (PIC#i), corresponding to respective identifiers (ID#i), provided, of said users, - implementing (S2) a recognition of the respective faces of the users in the scene, and - determining (S3) the current positions of the users in the scene corresponding to the respective identifiers, provided, of the users.

8. A method according to any one of claims 6 and 7, taken in combination with any one of claims 3 to 5, further comprising: - memorizing the identifier of the selected candidate as a possible interlocutor of the current speaker, corresponding to a given time, - estimating a score (S8), from a weighted average with a forgetting factor, over a sliding time window including the given time and one or more times preceding the given time, assigned to the memorized identifiers of the selected candidates, for each selected candidate at the given time and at the preceding times, and - determining (S10) the selected candidate having the best score as an interlocutor of the current speaker for the given time.

9. A method according to claim 8, wherein a weight greater than the weights of other retained candidates is assigned (S9) to a retained candidate who was a current speaker at at least one of said preceding moments.

10. A method according to any one of the preceding claims, wherein the scene is a videoconferencing room comprising a screen (ECR), and wherein it is determined (S5, ORI_k) that the possible interlocutor of the current speaker is in a remote site participating in the videoconference if the respective head orientations of the room users are detected as directed towards the screen, for a number of users greater than a threshold.

11. A method according to claim 10, wherein, if it is alternatively determined that the possible interlocutor of the current speaker is in the videoconference room, the users whose orientations Heads detected as being directed towards the screen are excluded from said user section.

12. A method according to any one of the preceding claims, comprising: - providing (S1) image data of user faces present in the scene, corresponding to respective identifiers, provided, of said users, - implementing (S2) a recognition of the respective faces of the users in the scene, - storing the identifier of the user determined as the interlocutor of the current speaker at a given time, for the purpose of transmitting (P7) said stored identifier, corresponding to the given time, to at least one script editor for displaying said stored identifier in verbatim transcripts of current speakers and interlocutors at respective given times.

13. A method according to any one of the preceding claims, wherein the scene is a videoconferencing room comprising equipment connected to at least one remote videoconferencing equipment, the method comprising: - detecting (P2; S2) the respective faces of the users in the scene, and - storing at least data of a position (POS_i) of the user determined as the interlocutor of the current speaker at a given time, for the purpose of transmitting (P7) the data of said position (POS_i), corresponding to the given time, to the remote videoconferencing equipment, configured to drive a display of the image of the scene at the given time with highlighting of the interlocutor.

14. A computer program comprising instructions for carrying out the method according to any one of the preceding claims when this program is executed by a processor.

15. Image data processing device for a scene comprising a plurality of users including a current speaker at a given time, comprising a processing circuit (MEM, PROC, COM) for implementing the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Group and conversational framing for speaker tracking in a video conference system

    US20180249124A1

  • Electronic apparatus and method for controlling thereof

    US20190320140A1

  • Information processing device, information processing method, and program

    US20200335105A1

  • Frame synchronous rendering of remote participant identities

    US20200344278A1

  • Video-conference endpoint

    WO2022248671A1