Participant reidentification in multi-camera videoconferencing including central camera
Patent Information
- Application Number
- EP2023814303
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2026-09-09
AI Technical Summary
Existing videoconferencing systems using multiple cameras face challenges in accurately identifying and reidentifying participants across different camera views, leading to duplication of participant images and suboptimal viewing experiences.
The implementation of an inside-out identification system that utilizes a central camera and a front camera, combined with an AI subject detector model, to determine participant locations and reidentify participants across multiple camera views, preventing duplicate images and selecting optimal views for transmission.
This approach effectively reduces participant duplication in videoconferencing systems, enhances the clarity of communication, and improves the overall videoconferencing experience by ensuring only optimal views of participants are transmitted.
Smart Images

Figure US2023036546_08052025_PF_FP_ABST
Abstract
Description
PARTICIPANT REIDENTIFICATION IN MULTI-CAMERA VIDEOCONFERENCING INCLUDING CENTRAL CAMERABACKGROUND
[0001] Videoconferencing systems typically connect people at a videoconferencing endpoint, such as a videoconference room, with people at other videoconferencing endpoints. In some videoconferencing modes, all participants detected in a videoconference room are separated and put into a gallery view to create equality with the remote participants.BRIEF DESCRIPTION OF THE DRAWINGS
[0002] FIG. 1 is a top view of an example conference room, according to some aspects of the present disclosure.
[0003] FIG. 2 is a schematic isometric view of another example conference room with three participants located at different coordinate positions in relation to a videoconference camera.
[0004] FIG. 3 is a top view of the example conference room of FIG. 2.
[0005] FIG. 4 is a schematic isometric view of yet another example conference room with two participants located at different coordinate positions, according to some examples of the present disclosure.
[0006] FIG. 5 is a schematic diagram of a camera and a two-dimensional image plane with an example determination of room coordinates for a head bounding box, according to an example of the present disclosure.
[0007] FIG. 6 is a front view of still another example conference room with seven individuals located at different coordinate positions in relation to a front camera and a central camera according to some examples of the present disclosure.
[0008] FIG. 7 is a forward-looking view of the example conference room of FIG. 6 captured by the central camera.
[0009] FIG. 8 is a left side-looking view of the example conference room of FIG. 6 captured by the central camera.
[0010] FIG. 9 is a rearward-looking view of the example conference room of FIG. 6 captured by the central camera.
[0011] FIG. 10 is a right side-looking view of the conference room of FIG. 6 captured by the central camera.
[0012] FIG. 11 is an example of an image strip including single views of individuals in the example conference room of FIG. 6, taken from the views captured by central camera of FIGS.
[0013] FIG. 12A is a pixel plot of an image of a single individual in the example conference room of FIG. 6 captured by the front camera.
[0014] FIG. 12B is a pixel plot of the individual of FIG. 12A after performing background segmentation.
[0015] FIG. 13 is a plot of Euclidean distances between embeddings associated with each individual in the image strip of FIG. 11, and embeddings associated with the individual of FIG. 12B, identified as participant two.
[0016] FIG. 14 is a front view of the example conference room of FIG. 6, where images of the individuals captured by the front camera are aligned with the images of the individuals captured by the central camera, as shown in the image strip of FIG. 11.
[0017] FIG. 15 is a flowchart of a method of implementing an inside-out identification system in a conference room, according to an example of the present disclosure.
[0018] FIG. 16 is a front view of a camera according to an example of the present disclosure.
[0019] FIG. 17 is a flowchart of a method of determining image plane coordinates for a detected subject, according to an example of the present disclosure.
[0020] FIG. 18 is a schematic of an example codec, according to an example of the present disclosure.DETAILED DESCRIPTION
[0021] In a video conference room equipped with multiple cameras, the same participant may appear in the field of view of more than one camera. Thus, an issue in videoconferencing using multiple cameras is the duplication of participants, which includes transmitting more than one image of a particular participant to a far end of the videoconference. Correspondingly, the framing of individuals in a videoconference room can be improved by determining the location of individual participants in the room relative to one another or a particular reference point. For example, if Person A is sitting at 2.5 meters from the camera and Person B is sitting at 4 meters from the camera, the ability to detect this location information can enable various advanced framing and tracking experiences. For example, participant location information can be used to compare multiple camera views and reidentify the same participant in each camera view to determine an optimal camera view of the participant.
[0022] More specifically, when multiple cameras of a videoconference system are used in a conference room with two or more participants, multiple views of the participants are captured, e. , multiple views at different angles. This is particularly true when the videoconference system includes a front camera and a central camera, such as a 360-degree camera, which areused to record visual data for the videoconference system. For example, a primary camera may capture a first or frontal view of the participants, and the central camera may capture views of the meeting participants from an inside-out perspective. As a result, multiple images of the same participant and / or participants may be transmitted to a far end of the videoconference, which in turn may be confusing to the other participants in the videoconference. Correspondingly, it is undesirable to transmit a particular view to a far end of the videoconference if a participant’s face is not fully visible in that particular view. Further, no adequate industry standard or specification has been developed to track individuals across multiple camera views in a videoconferencing system based on their relative distance from one another or their relative position in a conference room.
[0023] For many applications, it is useful to know the horizontal and vertical location of the participants in the room to provide for a more comprehensive and complete understanding of the videoconference room environment. The ability to determine two-dimensional room parameters, e.g. , a width and a height , for each meeting participant and use such parameters to identity’ each participant can be enabled by using an external feature prop or computationally intensive machine learning-based monocular depth estimation models, but such approaches impose significant hardware and / or processing costs without providing the accuracy for identifying participants across multiple camera views. Further, such approaches are limited in their adaptability to a variety of different room geometries and locations due to their reliance on dedicated hardware.
[0024] For example, various techniques attempt to perform what is called person reidentification (RelD) to match a participant captured by one camera to the same participant captured on another camera. In one variation, pre-trained deep learning RelD models are utilized to identify embeddings of a participant in a particular camera view and match those embeddings with embeddings found in other camera views. However, this technique relies upon resource-intensive computations that require considerable memory in order to be performed, and such computations still prove ineffective to differentiate between participants who have similar clothing to one another, such as participants who are wearing uniforms. In addition, deep learning RelD models are primarily trained by targeting upright pedestrians, so their performance is further compromised when such models are applied to seated participants.
[0025] In another variation, an RelD-based approach includes mapping three-dimensional (3D) world coordinate systems to two-dimensional (2D) camera planes by using an external object or feature vector matching to determine a transformation matrix. The reliance upon external reference points in a room limits the applicability of such models to the physical roomin which they are in. As a result, RelD techniques which rely on external reference points to calculate a transformation matrix lack adaptability and are unable to be used in a wide variety of videoconference room settings.
[0026] Accordingly, in some examples, the present disclosure provides apparatuses, methods, and media for reidentifying participants across multiple camera views in a videoconference. In particular, the present disclosure provides a method of identifying, from multiple camera views including a central camera view, each participant in a videoconferencing room, preventing duplicate participant images from being transmitted to a far end of the videoconference, and / or identify ing an optimal view of a participant to be transmitted to the far end instead of another, less desirable duplicate view. Thus, while a conference room is covered by multiple cameras, the far end site can be shown a single manipulated data stream that does not have duplicated people and that has the best view of everyone in the videoconferencing room. By utilizing the disclosed RelD methods, communication between participants in the videoconference may be clearer, and the overall videoconferencing experience may be more enjoyable for the participants. Further, the methods discussed herein are applicable to a wide variety of different locations and room designs, meaning that the disclosed methods may be easily assembled and applied to any particular conference room.
[0027] By way of example, FIG. 1 illustrates an example conference room 10 for use in videoconferencing. The conference room 10 includes a conference table 12 and a series of chairs 14. Persons 16 are seated in the chairs 14 around the conference table 12. In the nonlimiting example illustrated in FIG. 1, a first person 16A, a second person 16B, a third person 16C, and a fourth person 16D are seated around the conference table 12. While FIG. 1 illustrates an example of a videoconference room 10 with four persons 16, more or fewer persons 16 may be seated around the conference table 12 or otherwise situated within the conference room 10 at any given time. Further, more or fewer persons 16 may be located outside the conference room 10 at any given time. Additional examples of videoconference rooms, locations of participants in videoconference rooms, and camera arrangements will be discussed below in greater detail.
[0028] Referring still to FIG. 1, in some aspects, a videoconferencing system 18 can include cameras 20, microphone arrays 22, and a monitor 24. More specifically, as shown in the example of FIG. 1, the videoconferencing system 18 can include a first or front camera 20A and a second or central camera 20B. However, it is contemplated that the videoconferencing system 18 may include additional cameras (e.g., a third or left camera, a fourth or right camera, and / or other cameras). The front camera 20A has a field-of-view (FOV) 25, horizontal andvertical, and an axis or centerline (CL) 26 extending in a direction that corresponds to the direction in which the front camera 20A is pointing (z.e., the front camera’s 20A line of sight that is straight at 90-degrees from its focal point). For example, as shown in FIG. 1, the front camera 20A can have a horizontal FOV 25 A which pans horizontally, i.e., along a width dimension, in the conference room 10, and a vertical FOV 25B which pans vertically, i. e. , along a height dimension, in the conference room 10. The central camera 20B can be positioned on the conference table 12, e.g. , at a center of the conference table 12 or at another location on the conference table 12, and the central camera 20B can have a 360-degree FOV 25C. In some examples, the central camera 20B can be arranged as a camera array that includes multiple cameras, e.g., a front-facing camera, side-facing cameras, a rear facing camera, etc. Further, the front camera 20A can include a first microphone array 22A and the central camera 20B can include a second microphone array 22B and the microphone arrays 22 may be used to record and transmit audio data in the videoconference using sound source localization (SSL). In some examples, the first microphone array 22A is housed on or within a housing of the front camera 20A, and the second microphone array 22B is housed on or within a housing of the central camera 20B.
[0029] In addition, the videoconferencing system 18 can include a monitor 24 or television that is provided to display a far end conference site or sites and generally to provide loudspeaker output. The monitor 24 can be coupled to the front camera 20A and the first microphone array 22A, although it is contemplated that the monitor 24 can be positioned anywhere in the conference room 10, and that the videoconferencing system 18 may include additional monitors (not shown) positioned in the conference room 10.
[0030] Further, the centerline 26 of the front camera 20A is centered along the conference table 12. In some examples, the central camera 20B is positioned along the centerline 26 of the front camera 20A and on the conference table 12. In some aspects, a person 16 may be located within more than one of the FOVs 25 of the cameras 20, meaning that the person 16 may be duplicated when the captured views of the cameras 20 are transmitted to a far end site of the videoconference. This in turn may cause confusion in the videoconference and / or result in non- optimal views of the person 16 being transmitted to a far end of the videoconference. Thus, it is advantageous to reidentify the person 16 across each of the views captured by the cameras 20, thereby reducing confusion in the videoconference and ensuring that only one, optimal view of each person 16 is shown in the videoconference.
[0031] Correspondingly, the front camera 20A may provide a better view of the faces of certain persons 16 when they are looking forward, i. e. , toward the front of the conference room10, and the central camera 20B may provide a better view of the faces of certain persons 16 when they are facing a center of the table 12. In some aspects, the cameras 20 and microphone arrays 22 are used in combination to provide multiple views of the conference room 10, and the processes described herein can be applied to each of the views to reidentify participants thereacross and prevent duplicate images of the participants from being transmitted to a far end of the videoconference. For example, the systems, processes, and media described herein allow an optimal view of each meeting participant to be identified and chosen for transmission, while non-ideal views may not be transmitted to the far end of the videoconference. This is accomplished using an artificial intelligence (Al) or machine learning subject detector model, as discussed below. As used herein, an "optimal" view may be a view that provides a best frontal view of a participant, i. e.. a best view of a face of a participant compared to other views. For example, an optimal view may be determined by applying facial recognition techniques to an image. In particular, an optimal view may be a view in which the quality of facial features or embeddings recognized by a facial recognition technique is greater than any other view.
[0032] An example Al or machine learning human head detector model, which may also be referred to herein as a subject detector model, will now be described with reference to FIGS. 2-5. More specifically, referring now to FIGS. 2 and 3, another example conference room 40 is illustrated with three videoconference participants 42, 44, 46 located at different coordinate positions. In the conference room 40, the front camera 20A has a horizontal and vertical FOV, and the camera location with respect to the room 40 is denoted by the three-dimensional (3D) coordinates {0, 0, 0}. Further, the front camera 20A captures a view of all three participants 42, 44, 46 having locations that can be characterized in terms of a pan angle OPAN relative to a centerline 26 of the front camera 20A and a distance measure between the front camera 20A and each participant 42, 44, 46. In particular, a first participant 42 has a location defined by a first pan angle 48 and a first distance 50. In addition, a second participant 44 has a location defined by pan angle 52 and a second distance 54, and a third participant 44 has a location defined by pan angle 56 and a third distance measure 58.
[0033] Referring now specifically to FIG. 3, a top view is illustrated of the example conference room 40 of FIG. 2. In some examples, the location of each participant 42. 44, 46 may be characterized in terms of the pan angles 48, 52, 56 and distances 50, 54, 58 that are derived from an XROOM dimension or axis 60 and a yROOM dimension or axis 62, where the front camera 20A is located at {XROOM, VROOM} coordinate positions of {0, 0}. In particular, the first participant 42 has a location defined by the first pan angle 48 and a first distance measure 50 which is characterized by two-dimensional room distance parameters {-0.5, 1} to indicate thatthe participant is located at a “vertical” distance (in relation to the top view) of 1 meter, measured from the front camera 20A along the yROOM axis 62, and at a “horizontal” distance of -0.5 meters, measured along the XROOM axis 60 that is perpendicular to the VROOM axis 62. In addition, the second participant 44 has a location defined by the second pan angle 52 and a second distance measure 54 which is characterized by two-dimensional room distance parameters {0, 3} to indicate that the participant is located at a vertical distance of 3 meters (measured along the yRooM axis 62) and at a horizontal distance of 0 meters (measured along the XROOM axis 60) to indicate that the second person is located along the centerline 26 of the front camera 20 A, resulting in a zero-degree pan angle 52. Finally, the third participant 44 has a location defined by a third pan angle 56 and a third distance measure 58 which is characterized by two-dimensional room distance parameters {1. 2.5} to indicate that the participant is located at a vertical distance of 2.5 meters (measured along the VROOM axis 62) and at a horizontal distance of 1 meter (measured along the XROOM axis 60).
[0034] The relationship between the pan angle values (<DPAN) and the two-dimensional room distance parameters {XROOM, yROOM} may be determined by using a reference coordinate table (not shown) in which pan angle OP AN values for the videoconference front camera 20A are computed for meeting participants located at different coordinate positions {XROOM, VROOM} in the example conference room 40 of FIGS. 2 and 3. An identical table (not shown) of negative pan angle OP AN values (e.g, -OPAN) can be computed for coordinate positions of {-XROOM, yROOM} in the example conference room 40. Thus, it will be understood that the same pan angle d>PAN value (e.g., PAN = 0) will be generated for a meeting participant located along the centerline 26 of the front camera 20A (e.g., XROOM = 0) at any depth measure (e.g., yROOM = 0.5-8). Similarly, the same pan angle OP AN value (e.g., OP AN = 45) will be generated for a meeting participant located at any coordinate position where XROOM = yROOM. As illustrated, the pan angle OP AN alone may not be sufficient information for determining the two-dimensional room distance parameters {XROOM, yROOM} for the location of a participant. For example, the first participant 42 may appear larger to the front camera 20A than the second participant 44 due to vanishing points perspective. Thus, as a meeting participant moves further away from the front camera 20 A, the apparent height and width of the participant become smaller to the videoconferencing system, and when projected to a camera image sensor 64, meeting participants are represented with a smaller number of pixels compared to participants that are nearer to the front camera 20A. Further, if two heads are seen by the front camera 20 A as having the same size, they are not necessarily located at the same distance, and their locationsin a two-dimensional XROOM- OOM plane 66, as illustrated in FIG. 3, may be different due to the pan angle PAN and distortion in the height and width.
[0035] In particular, the statistical distribution of human head height and width measurements may be used to determine a min-median-max measure for the participant head size in centimeters. Additionally, by knowing the FOV resolution of the front camera 20A in both horizontal and vertical directions with the respective horizonal and vertical pixel counts, the measured angular extent of each head can be used to compute the percentage of the overall frame occupied by the head and the number of pixels for the head height and width measures. Using this information to compute a look-up table for min-median-max head sizes (height and width) at various distances, an artificial (Al) subject detector model can be applied to detect the location of each head in a two-dimensional viewing plane with specified image plane coordinates and associated width and height measures for a head frame or bounding box (e.g., {xbox, ybox, width, height}). By using the reverse look-up table operation, the distance can be determined between the front camera 20A and each head that is located on the centerline 26 of the front camera 20 A.
[0036] Referring specifically to FIG. 4, a front camera 20A is used to provide an image of a meeting participant taken along a two-dimensional image plane 110. The meeting participant can be located in a first, centered position 112 and a second, panned position 114 that is shifted laterally in the XROOM direction. In the first, centered position, the meeting participant is located along the centerline 26 of the front camera 20A (e.g. , OP AN = 0) at a distance, dO = Y meters, so the two-dimensional room distance parameters for the first, centered position 1 12 are {XROOM = 0, yROOM = Y}. In the second, panned position, the meeting participant is shifted laterally in the XROOM direction by a panned angle OP AN and is located at dl > dO meters, so the two- dimensional room distance parameters for the second, panned position 114 are {XROOM = P, yROOM = Y}. Further, the same vertical head height measure V / 2 for the meeting participant positions 112, 114 will result in an angular extent 0FRAME V I / 2 for the first meeting participant position 112 that is larger than the angular extent 0FRAME_V2 / 2 for the second meeting participant position 114. In effect, the fact that the second, panned position 114 is located further away from the front camera 20 A than the first, centered position 112 (dl > dO) results in the angular extent for the second, panned position 114 appearing to be smaller than the angular extent for the first, centered position 112 so that 0FRAME V I / 2 > 0FRAME V2 / 2.
[0037] From the foregoing, the issue is to find an angular extent for the entire head height 0HH and then represent it as a percentage of the full frame vertical field of view (VFrame_Percentage) which is then translated into the number of pixels the head will occupy (VHead Pixel Count) at a particular distance and at a pan angle OP AN. To this end, the angular extent for the entire head height 0mii for the first meeting participant location 112 may be calculated by starting with the equation, tan(0mii / 2) = (V / 2) / d0. Solving for the angular extent 01, the angular extent for the entire head height 0niii may be calculated as 0inii = 2 arctan((V / 2) / dO). In similar fashion, the angular extent for the entire head height 0HH2 for the second meeting participant position 114 located at the pan angle QPAN may be calculated by starting with the equation, tan(0nw2 / 2) = (V / 2) / dl, where dl = dO2+ P2. Solving for the angular extent 0mi2, the angular extent for the entire head height 0ini2 may be calculated as 0(iii2 = 2 x arctan((V / 2) / dl) = 2 x arctan((V / 2VdO2+ P2)). Based on this computation, the percentage of the frame occupied by the head height for the second meeting participant location 114 can be computed as VFrame_Percentage = 0ini2 / Vertical FOV. In addition, the corresponding number of pixels for the head height for the second meeting participant location 114 can be computed as VHead Pixel Count = VFrame Percentage x Vertical FOV in pixels. Based on the foregoing calculations, the angular extent for the entire head height 0HH = 0FRAME_V may be calculated at discrete distances of, for example, 0.5 meters in each of the XROOM and yROOM directions that are equivalent to various angular pan angles OP AN which may be listed in a look-up table (not shown).
[0038] FIG. 5 illustrates a front camera 20A and an example videoconference room 200 including a two-dimensional image plane 210 to illustrate how to calculate a vertical or depth room distance YROOM (meters) to the meeting participant location from the distance measure XROOM (meters) by calculating a direct distance measure HYP between the front camera 20A and the meeting participant location. The two-dimensional image plane 210 includes a plurality of two-dimensional coordinate points 212, 214, 216 that are defined with image plane 210 coordinates {xi, yi} as described above. In addition, a head bounding box 218 is defined with reference to the starting coordinate point {xi, yi} for the head bounding box 218, a Width dimension (measured along the xi axis), and a Height dimension (measured along the yi axis). To locate the vertical or depth room distance YROOM (meters) from the front camera 20A, a vertical angular extent (0) for the head bounding box 218 is computed as 0 = Height * V_FOV / V_PIXELS, where Height is the height of the head bounding box in pixels, where V_FOV is the Vertical FOV in degrees, and where V PIXELS is the Vertical FOV in Pixels. Next, avertical angular extent for the upper half of the head bounding box is computed (0 / 2) and used to derive the direct distance measure HYP between the front camera 20A and the meeting participant location, HYP = V_HEAD / (2 x tan(0 / 2)), where HYP is the direct distance measure to the meeting participant location at the pan angle OP AN. Finally, the vertical or depth room distance YROOM (meters) is derived from the direct distance measure HYP and the distance measure XROOM (meters) using Pythagorean’s Theorem, YROO XR00M2. In someexamples, the Width and Height dimensions of the head bounding box 218 can also be manipulated, e.g., scaled, to define an upper body bounding box 220 for the meeting participant. For example, the Width and Height dimensions of the head bounding box 218 can be expanded to define the upper body bounding box 220 which can surround the head bounding box 218, i.e., a meeting participant’s head, as well as a meeting participant’s upper body and clothing. This in turn can allow for more robust embeddings to be extracted from an image, as will be discussed below in greater detail.
[0039] With this understanding of the Al subject detector model, the present disclosure provides methods, devices, systems, and computer readable media to accurately detect and reidentify participants in a videoconferencing system using multiple cameras, e.g., a front camera and a central camera in an inside-out videoconferencing system. The location of each meeting participant or subject is determined by the Al subject detector model using room distance parameters, as discussed above. In particular, coordinates, e.g., image and / or world coordinates, are determined for each participant in each camera view. In some aspects, the world coordinates identified by the Al subject detector model are referred to as world coordinate points. Further, identification (ID) labels are assigned to each meeting participant detected by the Al subject detector model, and each identification label in a first view- captured by a first camera is grouped or paired with each identification label in a second view captured by a second camera. In some aspects, the identification labels in the second image are paired with the identification labels in the first image based on a distance between the identification labels in the first and second images. In another example, ID labels are assigned to each human head in each camera view based on their respective coordinates, e.g., angles and / or distances, from centroids that are calculated for each respective image using the coordinates for each participant.
[0040] In addition, a reference subject or anchor may be selected for each camera view, which may be known to be the same participant in each camera view . If this is not known, the primary anchor in a first image captured by a front camera can be analyzed for embeddingsand then compared to embeddings that are associated with each of the subjects detected in a second image captured by the central camera. Specifically, distances between the embeddings in each image can be measured, and a subject in the second image associated with a minimum distance relative to the other subjects can be matched with the primary anchor, i.e., selected as a secondary anchor for the second image. The ID labels may be aligned across all camera views by re-ranking the ID labels in a clockwise or counterclockwise order about each centroid, starting with the anchors in each image, to identify each participant across all camera views. Accordingly, it will be understood that multiple methods may be used to identify human heads across different camera views without departing from the scope of the present disclosure.
[0041] In view of the above, FIGS. 6-14 illustrate an example of an inside-out identification system and corresponding operation. In particular, FIGS. 6-10 illustrate yet another example conference room 300 which is defined by a front wall 302 (see FIG. 7), a left wall 304, a right wall 306, and a rear wall 308. A table 310 is located in the center of the conference room 300, and seven participants 312 are located within the conference room 300. For example, a first participant 312A, a second participant 312B, a third participant 312C. a fourth participant 312D, a fifth participant 312E, a sixth participant 312F, and a seventh participant 312G are located in the conference room 300. Specifically, the first, second, and third participants 312A, 312B, 312C are seated along a left side 314 of the table 310, the fourth participant 312D is seated along a rear side 316 of the table 310, and the fifth, sixth, and seventh participants 312E, 312F, 312G are seated along a right side 318 of the table 310. The inside-out identification system can include a first or front camera 320 (see FIG. 7) and a second or central camera 322 (see FIG. 6). In some aspects, the front camera 320 is coupled to a monitor 324, e.g, fastened on top of the monitor 324, and the monitor 324 is located at the front of the conference room 300 adjacent the front wall 302 of the conference room 300 (see FIG. 7). In some aspects, the central camera 322 is positioned on the table 310, e.g., a center of the table 310, and the central camera 322 may be an omnidirectional or 360-degree camera (see FIG. 6). The inside-out identification system can further include a processor and a memory coupled to the processor, and the memory stores programs instructions that, when executed by the processor, cause the processor to perform certain operations described herein, including performing videoconferencing operations such as transmitting data to a far end videoconferencing site. The processor and the memory7can be part of the front camera 320, the central camera 322, and / or a separate codec, as described below. Thus, some operations of the inside-out identification system can be performed within the front camera 320, the central camera 322. and / or the separate codec.
[0042] Referring specifically now to FIG. 6, a front or first image 326 is illustrated of the conference room 300, as captured by the front camera 320 (see FIG. 7). As illustrated, the left wall 304, the right wall 306, and the rear wall 308 are visible in the first image 326, and each of the participants 312 are also visible in the first image 326. Put another way, the walls 304, 306, 308, and the participants 312 are within a FOV of the front camera 320. After the first image 326 is captured, the Al subject detector model, as discussed above, can be applied to the first image 326 to determine world and / or image coordinates of each of the participants 312 and generate corresponding bounding boxes for the participants 312. In particular, the Al subject detector model can generate head bounding boxes 328A, 328B, 328C, 328D, 328E, 328F, 328G that correspond to the respective heads of the participants 312. The Al subject detector model can also manipulate the coordinates of the head bounding boxes 328 to further generate upper body bounding boxes 330A, 330B, 330C, 330D, 330E, 330F, 330G that correspond to the respective upper bodies of each of the participants 312. The upper body bounding boxes 330 may allow the upper body and / or clothing of each participant 312 to be considered when extracting features or embeddings from the first image 326, as will be discussed below in greater detail.
[0043] Once the coordinates are know n for each of the bounding boxes 328, 330, centers of each bounding box, e.g. , the head bounding boxes 328, the upper body bounding boxes 330, or both, are calculated and stored in the inside-out identification system. For example, a first centroid 332, as shown in FIG. 6. can be calculated using the centers of the bounding boxes 328, 330 using the below formulas:In the above formulas, xk,yk), k = 0,1, ... , n corresponds to the pixel coordinates of the center of each bounding box 328. 330 and (%c,yc) denotes the coordinates of the first centroid 332.
[0044] Once the first centroid 332 has been calculated, ID labels (not shown), e.g., 0, 1, 2 etc., may be assigned arbitrarily to each participant 312 by the Al subject detector model, or the ID labels may be assigned to each participant 312 based on a predetermined order. In one example, the ID labels may be assigned to each participant 312 based on the order in which the bounding boxes 328, 330 are calculated and generated, such as a clockwise or counterclockwise order starting with a participant closest to the front camera 320, e.g. , the first participant 312A. Thus, it will be understood that a variety' of methods may be used to assign the ID labels to the participant 312. Once the ID labels are assigned to the participants 312, the inside-out identification system measures the distance, e.g., a Euclidean distance, of each bounding box328, 330 center to the first centroid 332. In some aspects, the inside-out identification can rank each of the participants 312 in a clockwise or counterclockwise order about the first centroid 332, beginning with a participant 312 with the minimum Euclidean distance to the first centroid 332, e.g., the fourth participant 312D.
[0045] With continued reference to FIG. 6, the inside-out identification system can further identify a participant 312 as a reference person or anchor(s) in the first image 326 based on the position of the bounding boxes 328, 330 therein and / or image quality. In some aspects, an anchor is one of the participants 312 detected in the first image 326 that will undergo image processing, e.g., feature extraction and / or background segmentation, before being compared to the participants 312 detected by the central camera 322 to align the images captured by the front camera 320 and the central camera 322. In particular, an anchor can be a bounding box 328, 330 associated with any of the participants 312. In some aspects, an anchor is selected using the known coordinates of the participants 312. It is contemplated that the specific method of selecting an anchor may depend upon the geometry7of the room, meaning that the method of selecting an anchor may be modified to best suit a particular conference room. For example, the inside-out identification system may apply the equations in Table 1 below to the first image 326 to identify an anchor as the participant(s) closest to the front wall 302 (see FIG. 7).Table 1In the above table, argmin{xL. — yi) is used to identify the participant who is farthest toward a bottom-left comer 336 of the first image 326. and argmax xi +is used to identify the participant 312 who is farthest toward a bottom-right comer 338 of the first image 326, e.g., the seventh participant 312G. As a result, the inside-out identification system can identify7the front participants, e.g., the first and seventh participants 312A, 312F, as the anchors in the first image 326.
[0046] In addition, the inside-out identification system may select anchors based on an intersection-over-union (IOU) score associated with each participant 312 detected in the first image 326. To that end, the Al subject detection model may also calculate an IOU score for each of the participants 312 which can be a measure of the modeling accuracy of the bounding boxes 328, 330 to the actual heads and upper bodies, respectively, of the participants 312. After calculating IOU scores for each participant 312, the inside-out identification system may selectthe participant 312 with the lowest IOU score as an anchor, or the inside-out identification system may select any of the participants 312 with an IOU score less than about 50%. less than about 40%, less than about 25%, less than about 10%, or less than about 5% as anchors. In the non-limiting example illustrated in FIG. 6, the inside-out identification system may determine that the third and fourth participants 312C, 312D each have an IOU score of less than about 10%, so the inside-out identification system may subsequently select the third and fourth participants 312C, 312D as anchors.
[0047] Accordingly, according to this example, the inside-out identification system may select multiple participants 312, e.g., the first, third, fourth, and seventh participants 312A, 312C, 312D, 312E, as anchors to further undergo image processing. By selecting a subgroup of the participants 312 as anchors for downstream processing, the inside-out system can streamline subject reidentification, thereby reducing the amount of resources, e.g. , power, time, memory space, etc., needed to identify videoconference participants across different camera images. In some aspects, the inside-out identification system may identify one of the selected anchors, i.e., the first, third, fourth, and seventh participants 312A, 312C, 312D, 312E, as a primary anchor 340. For example, the primary anchor 340 may be selected based on the Euclidean distance of the participants 312 relative to the first centroid 332. In the non-limiting example illustrated in FIG. 6, the fourth participant 312D has the least Euclidean distance, i.e., is closest, to the first centroid 332, so the fourth participant 312D can be selected as the primary anchor 340. However, it is contemplated that selection of the primary anchor 340 can also be based on a variety of other factors, such as, for example, bounding box 328, 330 coordinate position, ID label ranking, pre-stored data associated with the participants 312, etc. Further aspects of the primary anchor 340 will be discussed below in greater detail.
[0048] FIGS. 7-10 illustrate additional images 342 of the conference room 300 that are captured by the central camera 322 (see FIG. 6). As discussed above, the central camera 322 may be configured as an omnidirectional camera and / or a camera array with multiple cameras to capture the additional images 342. For example, FIG. 7 illustrates a forward-looking or second image 342A of the conference room 300 captured by the central camera 322 positioned on the table 310. In the second image 342A, the first participant 312A and the seventh participant 312G are visible, along with the front camera 320 positioned along the front wall 302. In some aspects, the front camera 320 can be configured to rotate to focus on an active speaker, or the front camera 320 may be fixed in its position and does not rotate. As discussed above, the Al subject detector model can generate head bounding boxes 328A, 328G and upper body bounding boxes 330A, 330G for the first and seventh participants 312A, 312G,respectively in the second image 342A. It should also be noted that, according to some examples, the Al subject detector model can perform operations to generate bounding boxes 328, 330 natively within the particular camera 320, 322 that captures the respective views.
[0049] Correspondingly, FIG. 8 illustrates a left side-looking or third image 342B of the conference room 300 captured by the central camera 322 positioned on the table 310. In the third image 342B, the first, second, and third participants 312A. 312B, 312C are visible. As discussed above, the Al subject detector model can generate head bounding boxes 328 A, 328B, 328C and upper body bounding boxes 330A, 330B, 330C for the first, second, and third participants 312A, 312B, 312C, respectively, in the third image 342B.
[0050] In addition, FIG. 9 illustrates a rearw ard-looking view or fourth image 342C of the conference room 300 captured by the central camera 322 positioned on the table 310. In the fourth image 342C, the third, fourth, and fifth participants 312C, 312D, 312E are visible. As discussed above, the Al subject detector model can generate head bounding boxes 328C, 328D, 328E and upper body bounding boxes 330C, 330D, 330E for the third, fourth, and fifth participants 312C, 312D, 312E, respectively, in the fourth image 342C.
[0051] Further, FIG. 10 illustrates a right side-looking or fifth image 342D of the conference room 300 captured by the central camera 322 positioned on the table 310. In the fifth image 342D, the fifth, sixth, and seventh participants 312E, 312F, 312G are visible. As discussed above, the Al subject detector model can generate head bounding boxes 328 AE, 328F, 328G and upper body bounding boxes 330E, 330F, 330G for the fifth, sixth, and seventh participants 312E, 312F, 312G respectively, in the fifth image 342D.
[0052] It will be understood that FIGS. 7-10 illustrate examples of images captured by the central camera 322, and that the central camera 322 may capture more or few er images of the conference room 300 according to user preference and / or camera configuration. In some aspects, the inside-out identification system is configured to reidentify participants 312 in the additional images 342 to eliminate duplicate views of the participants 312 before the additional images 342 are compared to the first image 326. In particular, the inside-out system can compare embeddings in the additional images 342 to determine if any of the participants 312 have been duplicated, and then eliminate any duplications. For example, referring specifically to FIGS. 7 and 8, the second and third images 342A, 324B both include views of the first participant 312A. Accordingly, the inside-out identification system can extract embeddings associated with the first participant 312A, e.g., facial features, clothing, posture, etc., to determine that the right-most participant 312 in the second image 342A, i. e.. the first participant 312A, is the same person as the left-most participant 312 in the second image 342B, i.e., thefirst participant 312A. It is contemplated that this comparison process can be similarly applied to each of the additional images 342 to identify if any of the other participants 312 have been duplicated.
[0053] Still referring to FIGS. 7 and 8, the inside-out identification system can select a single view of the first participant 312A for downstream processing after the comparison process discussed above has been performed. In some aspects, a view of the first participant 312A is chosen if the first participant 312A is fully visible in that view and / or if a center of the first participant’s 312A head bounding box 328 A is closer to a center of the respective upper body bounding box 330A in that view than in another view. For example, the first participant 312A is fully visible in the second image 342A but may be only partially visible in the third image 342B. so the second image 342A may be chosen to show the first participant 312A. In addition, the center of the head bounding box 328A is closer to the center of the upper body bonding box 330A in the second image 342A relative to the same in the third image 342B, so the second image 342A may be chosen over the third image 342B to show the first participant 312A. As discussed above, it is contemplated that this view selection process can be similarly applied to each of the additional images 342 to select a single view of any duplicated participants 312.
[0054] Accordingly, the inside-out identification system can eliminate duplicate views of the participants 312 captured by the central camera 322 from being transmitted to a far end of the videoconference. In some aspects, the inside-out identification system can create or stitch together a strip of the selected views of the participants 312. Referring now to FIG. 11 , an example central camera strip 344 is illustrated that includes single views of each of the participants 312 that have been selected for downstream processing as a result of the comparison and view selection processes described above. In particular, the strip 344 can include views of the first and seventh participants 312A, 312G taken from the second image 342A (see FIG. 7), a view of the second participant 312B taken from the third image 342B (see FIG. 8), view s of the third and fourth participants 312C, 312D, taken from the fourth image 342C (see FIG. 9), and views of the fifth and sixth participants 312E, 312F taken from the fifth image 342D (see FIG. 10). In this way, only a single view of each participant 312 may be further processed downstream.
[0055] After duplicate view s of the participants 312 captured by the central camera 322 have been eliminated, the first image 326 (see FIG. 6) and the views in the strip 344 can be subjected to image processing techniques before being compared to reidentify the participants 312 across the different camera views. It is contemplated that a variety of image processingtechniques may be used, including background segmentation and feature or embedding extraction. To that end, background segmentation can be used to isolate the participants 312, e.g., the participants 312 heads and upper bodies, from the background, e.g, the conference room 300, the walls 302, 304, 306, 308, the table 310, etc. This, in turn, can allow the embedding extraction process as discussed below to generate stronger embeddings, thereby leading to more accurate and effective reidentification of the participants 312 across different views.
[0056] Referring now to FIG. 12A, an example pixel plot 400 is illustrated of a view of fourth participant 312D, i.e., the primary anchor 340, taken from the first image 326 (see FIG. 6). In particular, the pixel plot 400 includes the fourth participant’s 312D head 402 and upper body 404, in addition to the rear wall 308 and the table 310. In some aspects, the fourth participant 312D may be identified as a foreground 406 in the pixel plot 400, and the rear wall 308 and the table 310 may be collectively identified as a background 408 in the pixel plot 400. The inside-out identification system can apply a background segmentation or removal model to the pixel plot 400 to isolate the foreground 406 from the background 408, i. e. , remove the background 408. Relatedly, FIG. 12B illustrates another example pixel plot 410 that can be generated after the background segmentation model has been applied to the pixel plot 400 of FIG. 12A. Notably, the background 408 shown in the pixel plot 400 of FIG. 12A has been removed such that the foreground 406, i.e., the fourth participant 312D, is isolated in the pixel plot 410 of FIG. 12B. It is contemplated that the background segmentation model can be applied to all of the images 326 (see FIG. 6), 342 (see FIGS. 7-10) to isolate all of the detected participants 312, or the background segmentation model may only be applied to the primary anchor 340 in the first image 326 and the participants 312 in the strip 344 of FIG. 11.
[0057] As discussed above, embeddings associated with a particular participant 312 can be generated, extracted, and compared with subsequent associated with other participants to reidentify participants across different images. Embeddings can include, for example, facial features, clothing, posture, colors, etc. associated with a particular participant. In some aspects, an Al model can be used to generate embeddings for an image. After the embeddings have been generated, a variety of techniques e.g., K-means clustering and / or Euclidean distance determinations, may be used to classify the embeddings and compare clusters of embeddings to one another. In some examples, neural netw orks can also be used to generate and / or compare embeddings in different images, including, for example, deep-leaming neural networks, convolutional neural networks, feedforward neural networks, recurrent neural networks, radial basis function neural networks, etc.
[0058] Referring now to FIG. 13, an example chart 450 is illustrated comparing Euclidean distances measured between embeddings in a first image and embeddings in a second image. In particular, a first set of embeddings can be generated and extracted for the anchors detected in the first image 326 (see FIG. 6), or at least the primary anchor 340, i. e. , the fourth participant 312D (see FIG. 12B). In addition, second embeddings or second sets of embeddings can be generated and extracted for each of the participants 312 detected in the additional images 342 (see FIGS. 7-10), e.g, the participants 312 in the strip 344 of FIG. 11. After the embeddings have been extracted, the inside-out identification system can compare the first and second embeddings to determine which of the participants 312 in the strip 344 is most similar to the primary anchor 340. For example, Euclidean distances can be measured between the first embeddings associated with the primary anchor 340 and the second embeddings associated with each participant 312 in the strip 344, and the Euclidean distances can then be compared with one another.
[0059] In the non-limiting example illustrated in FIG. 13, the chart 450 includes an x-axis 452 that can be an index of the participants 312 in the strip 344 (i.e., participant 0, 1, 2. etc.), a y-axis 454 that can define a range of Euclidean distances, and a data line 456 to show each of the measured Euclidean distances. In some aspects, Euclidean distance can be inversely proportional to a measure of similarity' between the primary' anchor 340, such as shown in FIG. 12B, and the participants 312 in the strip 344 shown in FIG. 11. In this way, a participant 312 in the strip 344 with the minimum Euclidean distance relative to the primary anchor 340 can be selected as a secondary' anchor 458. For example, the fourth participant 312D in the strip 344 can return a minimum Euclidean distance as indicated by circle 460, meaning that the fourth participant 312D in the strip 344 is most similar to the primary anchor 340 (see FIG. 12B). Accordingly, the inside-out identification system can select the fourth participant 312D in the strip 344 as the secondary anchor 458. After the secondary anchor 458 has been identified, the participants 312 can be rearranged beginning with the primary anchor 340 and the secondary' anchor 458, respectfully, to generate aligned front and central camera views of the participants 312. In some aspects, this process is repeated for each anchor that is identified in the first image 326, meaning that multiple secondary anchors may be selected as a result of the Euclidean distance comparison between the embeddings in the first image 326 and the embeddings in the strip 344.
[0060] In some examples, the inside-out identification system may be configured to use a multi-anchor scheme to confirm that the secondary anchor 458 has been correctly matched to the primary anchor 340. To that end, the inside-out identification system may initially identifythe primary anchor 340, as described above, as well as additional primary anchors 340, such as three additional primary anchors 340 from the first image. A chart 450. like that shown in FIG. 13, can be produced comparing Euclidean distances measured between embeddings in the first image and embeddings in the second image. A participant 312 in the strip 344 with the minimum Euclidean distance relative to the primary anchor 340 can be selected as a potential secondary anchor 458, as described above. Furthermore, two more participants 312 in the strip 344 with the next minimum Euclidean distances can be selected as potential secondary anchors 458.
[0061] According to this example, referring to FIG. 13, the first participant 312A, the fourth participant 312D, and the seventh participant 312G, in the strip 344 with the three least Euclidean distances relative to the primary anchor 340 can be considered potential matches or secondary anchors 458 to the primary anchor 340. The participants 312 may then be rearranged in three different orders, i.e., with each order beginning with one of the three potential secondary anchors 458, and these orders can then be matched to, e.g, aligned with, the order of the participants 312 in the first image 326 starting with the primary anchor 340, for example in a table like table 462 illustrated in FIG. 14. For each of these three orders (e.g., for each respective table 462), Euclidean distances can be measured between each of the three additional primary anchors 340 from the first image 326 and their supposed matched participant 312 in the strip 344. and then the three distances are summed along with the first Euclidean distance between the primary anchor 340 and the respective potential secondary anchor 458. This results in a Euclidean distance summation for each potential secondary anchor 458 as the sum of those four distances. The inside-out identification can then compare the three Euclidean distance summations to identify the minimum thereof, and the order of the participants 312 corresponding to the minimum Euclidean distance summation can be chosen as the correct order. For example, based on this multi-anchor method, the inside-out identification system may identify the order beginning with the fourth participant 312D as the order with the minimum Euclidean distance summation, meaning that the fourth participant 312D can be chosen as the true secondary anchor 458. Accordingly, it will be understood that the inside-out identification system can generate aligned views of the participants 312 can be using a variety of different methods.
[0062] FIG. 14 illustrates the first image 326 of the conference room 300 and a table 462 of aligned views of each of the participants 312. As discussed above, the participants 312 captured by the front camera 320 and the central camera 322 can be re-arranged beginning with the primary anchor 340 and the true secondary anchor 458, respectfully, i.e., the fourthparticipant 312D. As an example, images of each of the participants 312 captured in the first image 326 (see FIG. 6) and the strip 344 (see FIG. 11) are arranged in the table 462 above the conference room 300. A first row 464 of the table 462 can include images of each of the participants 312 captured by the central camera 322, and a second row 466 of the table 462 can include images of the participants captured by the front camera 320. The first row 464 begins with the secondary anchor 458, and the second row 466 begins with the primary anchor 340. The remaining participants 312 are then arranged in the rows 464, 466 in a clockwise order about the first centroid 332, i.e., the order in which they are seated around the table 310.
[0063] Thus, it will be understood that the order of the participants 312 can be aligned across different images by ranking the participants 312 in a clockwise order beginning with an anchor, e.g., the primary anchor 340 and the secondary anchor 458. Once the rankings are aligned, the inside-out identification system can reassign the ID labels in the table 462 to reflect the aligned order of the participants 312. Put another way, the ID labels associated with the participants 312 can be aligned across each of the images, e.g., the first image 326 of FIG. 6 and the additional images 342 of FIGS. 7-10. Accordingly, while the fourth participant 312D, the secondary anchor 458, was initially assigned with a third ID label (label 2) in the strip 344, as shown in FIG. 13, when re-aligned in the table 462, the fourth participant 312D is assigned a first ID label (label 0), as shown in FIG. 14. Similarly, while the second and third participants 312B, 312C were initially assigned with first and second ID labels (labels 0. 1), respectively, in the strip, as shown in FIG. 13. when re-aligned using a clockwise order in the table 462. the second and third participants 312B, 312C are shifted to the end of the row 464 and assigned with sixth and seventh ID labels (labels 5, 6), respectively, as shown in FIG. 14.
[0064] Therefore, the inside-out identification system disclosed herein is capable of reidentifying videoconference participants across different camera views. Correspondingly, the inside-out identification system prevents multiple views of the same participant from being transmitted to a far end of a videoconference, which in turn may reduce confusion in the videoconference. In some aspects, the inside-out identification system disclosed herein is particularly advantageous in crowded conference rooms, when there is little distance between participants in a conference room, and / or conference rooms including a front camera and a central camera. Further, it is contemplated that FIGS. 6-14 illustrate non-limiting examples of the inside-out identification system, and that the inside-out identification system may be applied to a variety of different conference rooms and is compatible with a variety of different camera arrangements. For example, in some applications, the inside-out identification system may incorporate a front camera, a central camera, as well as additional outside-in viewingcameras, such as rear camera, left-side camera, and / or right-side camera, etc., in which views captured by those additional outside-in cameras may be processed in a similar manner as described above.
[0065] FIG. 15 illustrates a method 500 of implementing the inside-out identification system discussed above. At step 502 images of a location are captured using a first or front camera and a second or central camera (or cameras). As discussed above, the front camera can be arranged at a front of a conference room, and the central camera(s) may be arranged on a table in the conference room and / or configured as an omnidirectional camera. In some aspects, the front camera, the central camera, or both are in communication with and / or connected to a monitor and / or a codec that includes a memory and a processor, as will be discussed below in greater detail. At step 504. human heads in the images are detected using an Al subject detection model, as described above. For example, the Al subject detection model is applied to images captured by each of the cameras in order to identify, for each detected human head, a head bounding box with specified room and / or pixel coordinates. The Al subject detection model also determines coordinates for each bounding box, and may also extract embeddings associated with each detected human head. Step 504 can also include detecting body bounding boxes in some applications, as described above. At step 506, a centroid is determined based on the coordinates of the bounding boxes, or, more specifically, based on the coordinates of a center of each bounding box. At step 508, a first participant in the image captured by the front camera is identified as a primary anchor. The primary anchor may be selected based on image quality and / or bounding box Coordinates. At step 510, Euclidean distances are measured between embeddings in the different images, i. e. , the images captured by the front camera and the central camera, to match the primary anchor identified in the image captured by the front camera to a secondary anchor in an image capture by the central camera. In some aspects, the secondary anchor is a participant having embeddings detected in the image captured by the central camera which are at a minimum Euclidean distance to the embeddings of the primary anchor relative to the other subjects in the image captured by the central camera. At step 512, the ID labels are aligned in counterclockwise order, for example, starting with a first ID label associated with the bounding box that has a minimum polar angle with respect to the centroid. In some aspects, this ranking is stored in the inside-out identification system. In this way, the inside-out identification system correlates the ID labels across the images captured by the front and central cameras, which in turn allows the system to track participants across the images. In some aspects, the inside-out identification system repeats each step in the method 500 during normal operation, meaning that the centroid-based identification reidentifies participantscontinuously as the first and second cameras capture images of the location. Generally, the method 500 can be performed in real-time or near real-time. For example, in some aspects, the steps 502, 504, 506, 508, 510, 512, of the method 500 are repeated after a period of time has elapsed, such as, e.g., at least every 30 seconds, or at least every 15 seconds, or at least every 10 seconds, or at least every 5 seconds, or at least every 3 seconds, or at least every second, or at least every 0.5 seconds.
[0066] It should be noted that the above method 500, or any methods or processes described herein, can be implemented as a set of instructions, tangibly embodied on a non-transitory computer-readable media, such that a processor device can implement the instructions based upon reading the instructions from the computer-readable media.
[0067] FIG. 16 illustrates an example camera 620, which may be similar to the front camera 20A shown in FIG. 1, or front camera 320 shown in FIG. 7, and an example microphone array 622, similar to the first microphone array 22A show n in FIG. 1. The camera 620 has a housing 624 with a lens 626 provided in the center to operate with an imager 628. A series of microphone openings 630. such as five openings 630. are provided as ports to microphones in the microphone array 622. In some examples, the openings 630 form a horizontal line 632 to provide a desired angular determination for the SSL process, as discussed above. FIG. 16 is an example illustration of a camera 620, though numerous other configurations are possible, with varying camera lens and microphone configurations. Additionally, in some examples, aspects of the technology, including computerized implementations of methods according to the technology, can be implemented as a system, method, apparatus, or article of manufacture using standard programming or engineering techniques to produce software, firmware, hardware, machine readable instructions, or any combination thereof to control a processor device (e.g., a serial or parallel general purpose or specialized processor chip, a single- or multicore chip, a microprocessor, a field programmable gate array, any variety of combinations of a control unit, arithmetic logic unit, and processor register, and so on), a computer (e.g, a processor device operatively coupled to a memory), or another electronically operated controller to implement aspects detailed herein. Accordingly, for example, the technology can be implemented as a set of instructions, tangibly embodied on a non-transitory computer- readable media, such that a processor device can implement the instructions based upon reading the instructions from the computer-readable media. Some examples of the technology7can include (or utilize) a control device such as, e.g., an automation device, a special purpose or general-purpose computer including various computer hardware, software, firmware, and so on, consistent with the discussion below7. As specific examples, a control device can include aprocessor, a microcontroller, a field-programmable gate array, a programmable logic controller, logic gates etc., and other suitable components for implementation of appropriate functionality (e.g., memory, communication systems, power sources, user interfaces and other inputs, etc.).
[0068] As described above, the methods of some aspects of the present disclosure include detecting a location of individual meeting participants using an Al subject detector model. Referring now to FIG. 17, an example process 700 is illustrated for determining coordinates of a detected human head using such an Al subject detector process and, in particular, an Al human head detector process. The Al human head detector process analyzes incoming roomview video frame images 702 of a meeting room scene with a machine-learning, Al human head detector model 704 to detect and display human heads with corresponding head bounding boxes 706, 708, 710. In some examples, width and height dimensions of the head bounding boxes 706, 708, 710 can be manipulated, e.g., to define upper body bounding boxes 712. As depicted, each incoming room- view video frame image 702 may be captured by a front camera 20A in the video conferencing system. Each incoming room-view video frame image 702 may be processed with an on-device Al human head detector model 704 that may be located at the respective camera which captures the video frame images. However, in other examples, the Al human head detector model 704 may be located at a remote or centralized location, or at only a single camera. Wherever located, the Al human head detector model 704 may include a plurality of processing modules 714. 716, 718. 720 which implement a machine learning model that is trained to detect or classify human heads from the incoming video frame images, and to identify, for each detected human head, a head bounding box with specified image plane coordinate and dimension information.
[0069] In this example, the Al human head detector model 704 may include a first preprocessing module 714 that applies image pre-processing (such as color conversion, image scaling, image enhancement, image resizing, etc.) so that the input video frame image is prepared for subsequent Al processing. In addition, a second module 716 may include training data parameters and / or model architecture definitions which may be pre-defined and used to train and define the human head detection model 704 to accurately detect or classify human heads from the incoming video frame images. In selected examples, a human head detection model module 718 may be implemented as a model inference software or machine learning model, such as a Convolutional Neural Network (CNN) model that is specially trained for video codec operations to detect heads in an input image by generating pixel-wise locations for each detected head and by generating, for each detected head, a corresponding head bounding boxwhich frames the detected head. Finally, the Al human head detector model 704 may include a post-processing module 720 which is applies image post-processing to the output from the Al human head detector model module 718 to make the processed images suitable for human viewing and understanding. In addition, the post-processing module 720 may also reduce the size of the data outputs generated by the human head detection model module 718, such as by consolidating or grouping a plurality of head bounding boxes or frames which are generated from a single meeting participant so that a single head bounding box or frame is specified.
[0070] Based on the results of the processing modules 716, 718, 720, the Al human head detector model 704 may generate output video frame images 702 in which the detected human heads are framed with corresponding head bounding boxes 706, 708, 710 and the detected upper bodies are framed with corresponding upper body boxes 712. As depicted, the first output video frame image 702a includes head bounding boxes 706a, 706b, and 706c and upper body bounding boxes 712a, 712b, 712c which are superimposed around each detected human head and upper body, respectively. In addition, the second output video frame image 702b includes head bounding boxes 708a, 708b, and 708c and upper body bounding boxes 712a, 712b, 712c which are superimposed around each detected human head, and upper body respectively. Further, the third output video frame image 702c includes head bounding boxes 710a, 710b and upper body bounding boxes 712a, 712b which are superimposed around each detected human head and upper body, respectively. The Al human head detector model 704 may specify each head bounding box using any suitable pixel-based parameters, such as defining the x and y pixel coordinates of a head bounding box or frame in combination with the height and width dimensions of the head bounding box or frame. In addition, the Al human head detector model 704 may specify a distance measure between the camera location and the location of the detected human head using any suitable measurement technique. The Al human head detector model 704 may also compute, for each head bounding box, a corresponding confidence measure or score which quantifies the model’s confidence that a human head is detected.
[0071] In some examples of the present disclosure, the Al human head detector model 704 may specify all head detections in a data structure that holds the coordinates of each detected human head along with their detection confidence. More specifically, the human head data structure for a number, n, of human heads may be generated as follows:In this example, xi and yi refer to the image plane coordinates of the ithdetected head, and where Widthi and Heights refer to the width and height information for the head bounding box of the ithdetected head. In addition, Scores is in the range [0, 100] and reflects confidence as a percentage for the ilhdetected head. This data structure may be used as an input to various applications, such as framing, tracking, composing, recording, switching, reporting, encoding, etc. In this example data structure, the first detected head is in the image frame in a head bounding box located at pixel location parameters xi, yi and extending laterally by Widthi and vertically down by Heighti. In addition, the second detected head is in the image frame in a head bounding box located at pixel location parameters X2, y and extending laterally by Widthr and vertically down by Height2, and the nthdetected head is in the image frame in a head bounding box located at pixel location parameters xn, yn and extending laterally by Widthn and vertically down by Heightn. In some aspects, the center of each head bounding box isdetermined using the following equation: Oh + ~ , y(+ ~) ■
[0072] This human head data structure may then be used as an input to the distance estimation process that takes the {Width, Height} parameters of each head bounding box to pick the best matching distance in terms of meeting room coordinates {XROOM, yROOM} from the look-up table (described above) by first using one of the Width or Height parameters with a first look-up table, and then using the other parameter as a tie breaking if multiple meeting room coordinates {XROOM, YROOM} are determined using the first parameter. The human head data structure itself may then be modified to also embed the distance information with each Head, resulting in a modified human head data structure that looks like the following:where {XROOMI, VROOMI}, {XROOMZ, yROOM2}, ... , {XROOMH, VROOMH} specify the distance of Headi, Head2, . . . , Headn, from the camera, respective, in two-dimensional coordinates.
[0073] FIG. 18 illustrates aspects of a codec 800 according to some examples of the present disclosure. As discussed above, a codec 800 may be a separate device of a videoconferencing system or may be incorporated into the camera(s) within the videoconferencing system, such as a primary camera. Generally, the codec 800 includes machine readable instructions to maintain a video call with a videoconferencing end point, receive streams from secondary cameras (and a primary camera if not integrated with the primary camera), and encode and composite the streams, according to the methods described herein, to send to the end point.
[0074] As shown in FIG. 18, the codec 800 may include loudspeaker(s) 802, though in many cases the loudspeaker 802 is provided in the monitor 804. The codec 800 may include microphone(s) 806 interfaced via a bus 808. The microphones 806 are connected through an analog to digital (AID) converter 810, and the loudspeaker 802 is connected through a digital to analog (D / A) converter 812. The codec 800 also includes a processing unit 814, a network interface 816, a flash or other non-transitory memory’ 818, RAM 820, and an input / output (I / O) general interface 822, all coupled by a bus 808. A camera 824 is connected to the I / O general interface 822. Microphone(s) 806 are connected to the network interface 816. An HD MI interface 826 is connected to the bus 808 and to the external display or monitor 804. Bus 808 is illustrative and any interconnect between the elements can used, such as Peripheral Compo-nent Interconnect Express (PCie) links and switches, Universal Serial Bus (USB) links and hubs, and combinations thereof. The camera 824 and microphones 806, 806 can be contained in housings containing the other components or can be external and removable, connected by wired or wireless connections.
[0075] The processing unit 814 can include digital signal processors (DSPs), central processing units (CPUs), graphics processing units (GPUs), dedicated hardware elements, such as neural network accelerators and hardware codecs.
[0076] The flash memory' 818 stores modules of vary ing functionality’ in the form of software and firmware, generically programs or machine readable instructions, for controlling the codec 800. Illustrated modules include a video codec 828, camera control 830, framing 832, other video processing 834, audio codec 836, audio processing 838, network operations 840, user interface 842 and operating system, and various other modules 844. In some examples, an Al subject detector module is included with the modules included in the flash memory 818. Furthermore, in some examples, machine readable instructions can be stored in the flash memory 818 that cause the processing unit 814 to carry out any of the methods described above. The RAM 820 is used for storing any of the modules in the flash memory 818 when the module is executing, storing video images of video streams and audio samples of audio streams and can be used for scratchpad operation of the processing unit 814.
[0077] The network interface 816 enables communications between the codec 800 and other devices and can be wired, wireless or a combination. In one example, the network interface 816 is connected or coupled to the Internet 846 to communicate with remote endpoints 848 in a videoconference. In one example, the general interface 822 provides data transmission with local devices (not shown) such as a keyboard, mouse, printer, projector, display, exter-nal loudspeakers, additional cameras, and microphone pods, etc.
[0078] In one example, the camera 824 and the microphones 806 capture video and audio, respectively, in the videoconference environment and produce video and audio streams or signals transmitted through the bus 808 to the processing unit 814. As discussed herein, capturing “views” or “images” of a location may include capturing individual frames and / or frames within a video stream. For example, the camera 824 may be instructed to continuously capture a particular view, e.g. , images within a video stream, of a location for the duration of a videoconference. In one example of this disclosure, the processing unit 814 processes the video and audio using processes in the modules stored in the flash memory 818. Processed audio and video streams can be sent to and received from remote devices coupled to network interface 816 and devices coupled to general interface 822.
[0079] Microphones in the microphone array used for SSL can be used as the microphones providing speech to the far site, or separate microphones, such as microphone 806, can be used.
[0080] Certain operations of methods according to the technology7, or of systems executing those methods, can be represented schematically in the figures or otherwise discussed herein. Unless otherwise specified or limited, representation in the figures of particular operations in particular spatial order can not necessarily require those operations to be executed in a particular sequence corresponding to the particular spatial order. Correspondingly, certain operations represented in the figures, or otherwise disclosed herein, can be executed in different orders than are expressly illustrated or described, as appropriate for particular examples of the technology. Further, in some examples, certain operations can be executed in parallel, including by dedicated parallel processing devices, or separate computing devices that interoperate as part of a large system.
[0081] The disclosed technology is not limited in its application to the details of construction and the arrangement of components set forth in the following description or illustrated in the following drawings. Other examples of the disclosed technology are possible and examples described and / or illustrated here are capable of being practiced or of being carried out in various ways.
[0082] A plurality of hardware and software-based devices, as well as a plurality of different structural components can be used to implement the disclosed technology. In addition, examples of the disclosed technology can include hardware, software, and electronic components or modules that, for purposes of discussion, can be illustrated and described as if the majority of the components were implemented solely in hardware. However, in one example, the electronic based aspects of the disclosed technology can be implemented in software (for example, stored on non-transitory computer-readable medium) executable by aprocessor. Although certain drawings illustrate hardware and software located within particular devices, these depictions are for illustrative purposes. In some examples, the illustrated components can be combined or divided into separate software, firmware, hardware, or combinations thereof. As one example, instead of being located within and performed by a single electronic processor, logic and processing can be distributed among multiple electronic processors. Regardless of how they are combined or divided, hardware and software components can be located on the same computing device or can be distributed among different computing devices connected by a network or other suitable communication links.
[0083] Any suitable non-transitory computer usable or computer readable medium may be utilized. The computer-usable or computer-readable medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium would include the following: a portable computer diskette, a hard disk, a randomaccess memory (RAM), a read-only memory' (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, or a magnetic storage device. In the context of this disclosure, a computer-usable or computer-readable medium may be any medium that can contain, store, communicate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
[0084] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms ‘'component,” “system,” “module,” “block,” and the like are intended to encompass part or all of computer-related systems that include hardware, software, a combination of hardware and software, or software in execution. For example, a component can be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a computer and the computer can be a component. Components (or system, module, and so on) can reside within a process or thread of execution, can be localized on one computer, can be distributed between two or more computers or other processor devices, or can be included within another component (or system, module, and so on).
Claims
CLAIMS1. A method of identifying participants in a location using a front camera and a central camera, the method comprising: capturing images of the location using the front camera and the central camera; applying a machine learning subject detector model to the images to identify coordinates and extract embeddings for each participant detected in the images; determining a centroid of the coordinates for each participant detected in the images; applying, in each image, identification labels to each participant; identifying a first participant in a first image captured by the front camera as a primary anchor; measuring distances between the embeddings in the images to match the primary anchor to a secondary anchor detected in a second image captured by the central camera; and aligning the identification labels associated with each participant across each of the images based on the match.
2. The method of claim 1, wherein the central camera is configured as a 360- degree camera.
3. The method of claim 1, wherein the applying the machine learning subject detector model further includes generating head bounding boxes and upper body bounding boxes for each participant detected in the images.
4. The method of claim 1, further comprising applying a background removal model to the images to isolate each participant’s upper body and head from background in the images.
5. The method of claim 1, wherein identifying the first participant as the primary anchor includes one of identifying the first participant having less than 10% intersection over union (IOU) score, and identifying the first participant that is to the centroid.
6. The method of claim 1, wherein measuring distances between the embeddings in the images includes: comparing Euclidean distances between first embeddings associated with the primary anchor and second embeddings associated with each participant detected in the second image; andselecting a second participant detected in the second image with a minimum Euclidean distance between the first embeddings and the second embeddings as the secondary anchor.
7. The method of claim 6, wherein rearranging the identification labels further includes rearranging the identification labels in the second image in clockwise order staring with the secondary anchor.
8. A system for identifying participants in a location, the system comprising: a front camera to capture a first image of the location; a central camera to capture a second image of the location; a processor connected to the front camera, the central camera, or both, the processor to execute a program to perform videoconferencing operations including transmitting data to a far end videoconferencing site; and a memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to: identify coordinates and extracting embeddings for each participant detected in the first image and the second image; determine a centroid of the coordinates for each participant detected in the first image and the second image; apply, in the first image and the second image, identification labels to each participant; identify a first participant in the first image as a primary anchor; compare the embeddings in the first image and the second image to match the primary anchor to a secondary anchor detected in the second image; and align the identification labels associated with each participant across the first image and the second image based on the match.
9. The system of claim 8, wherein the central camera is configured as a 360-degree camera.
10. The system of claim 8, wherein a machine learning subject detector model is applied to the first image and the second image to generate head bounding boxes and upper body bounding boxes for each participant.
11. The system of claim 8, wherein a background removal model is applied to the first image and the second image to isolate each participant’s upper body and head from background.
12. The system of claim 8, wherein the processor is to identify the primary anchor as a participant that has less than 10% IOU, or a closest participant to the centroid.
13. The system of claim 8, wherein Euclidean distances are measured between first embeddings associated with the primary anchor and second embeddings associated with each participant detected in the second image, and wherein a second participant detected in the second image with a minimum Euclidean distance between the first embeddings and the second embeddings is selected as the secondary anchor.
14. A non-transitory computer-readable medium containing instructions that when executed cause a processor to: instruct a front camera to capture a first image of a location and a central camera to capture a second image of the location; apply a machine learning subject detector model to the first image and the second image to identify coordinates and extract embeddings for each participant detected in the first image and the second image; apply, in the first image and the second image, identification labels to each participant; identify a first participant in the first image as a primary anchor; compare the embeddings in the first image and the embeddings in the second image to match the primary anchor to a secondary anchor detected in the second image; and align identification labels associated with each participant across the first image and the second image based on the match.
15. The non-transitory computer-readable medium of claim 14, wherein Euclidean distances are measured between first embeddings associated with the primary anchor and second embeddings associated with each participant detected in the second image, and wherein a second participant detected in the second image with a minimum Euclidean distance between the first embeddings and the second embeddings is selected as the secondary anchor.