Video Conferencing Endpoints

JP2024521292A5Pending Publication Date: 2025-05-23NEATFRAME LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023566604
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-28
Filing Date
2022-05-27
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing video conferencing systems struggle to reliably detect and frame individuals within a spatial boundary, particularly in environments with glass walls or open spaces, leading to unwanted persons being included in the video stream.

Method used

A computer-implemented method and endpoint that define spatial boundaries using distance and angle criteria, identify individuals within the field of view, estimate their positions, and generate cropped video signals for transmission, ensuring only those within the defined boundaries are framed.

Benefits of technology

Enhances the reliability of framing individuals by accurately excluding those outside the specified spatial boundaries, improving the quality of video conferencing by focusing on intended participants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented method of operating a videoconferencing endpoint, the videoconferencing endpoint including a video camera that captures images indicative of a field of view, the method including the steps of receiving data defining spatial boundaries within the field of view, the spatial boundaries being defined at least in part by distance from the video camera, capturing images of the field of view, identifying one or more people within the field of view of the video camera, estimating a position of the or each person within the field of view of the video camera, and generating one or more video signals for transmission to a receiver that include one or more crop regions corresponding to the one or more people determined to be within the spatial boundaries.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a computer-implemented method and a videoconferencing endpoint. [Background technology]

[0002] In recent years, video conferencing and video calling have gained great popularity, allowing multiple users in different locations to have face-to-face discussions without the need to travel to the same location. Business meetings, remote lessons with students, and private video calls between friends and family are common applications of video conferencing technology. Video conferencing may be conducted using a smartphone or tablet, by a desktop computer, or by a dedicated video conferencing device (sometimes referred to as an endpoint).

[0003] A videoconferencing system allows both video and audio to be transmitted over a digital network between two or more participants located at different locations. Video input can be provided by a video camera or webcam located at each of the different locations, and audio input can be provided by a microphone at each of the different locations. Video output can be provided by a screen, display, monitor, television, or projector at each of the different locations, and audio output can be provided by speakers at each of the different locations. Hardware or software-based encoder-decoder technology compresses the analog video and audio data into digital packets for transmission over the digital network and decompresses the data for output at the different locations.

[0004] Some video conferencing systems include automatic framing algorithms that find and frame people in a conference room, e.g., by separating them from the existing video stream, cropping an area that includes all of them, or presenting them as individual video streams. In some cases, e.g., in rooms with glass walls or doors, or in open spaces, unwanted people outside the call (i.e., not participating in the call) may be detected and framed. Thus, it is desirable to improve the reliability with which people in a video call are detected and framed. Summary of the Invention

[0005] Accordingly, in a first aspect, an embodiment of the present invention provides a computer-implemented method of operating a videoconferencing endpoint, the videoconferencing endpoint including a video camera that captures images indicative of a field of view, the method comprising: receiving data defining a spatial boundary within a field of view, the spatial boundary being defined at least in part by a distance from a video camera; taking an image of the field of view; identifying one or more people within a field of view of a video camera; estimating the position of the or each person within the field of view of a video camera; generating one or more video signals including one or more crop regions corresponding to one or more people determined to be within the spatial boundary for transmission to a receiver; The present invention provides a computer-implemented method, comprising:

[0006] By defining spatial boundaries and framing only those determined to be within the boundaries, the reliability with which people are framed during a video call is increased.

[0007] We now describe optional features of the present invention, which may be applied alone or in any combination with any aspect of the present invention.

[0008] Generating the one or more video signals may include determining from the estimated position(s) that at least one of the one or more people is within the spatial boundary, and framing the one or more people determined to be within the spatial boundary to generate respective crop regions. Generating the one or more video signals may include framing the one or more people within a field of view of the camera to generate one or more crop regions, determining from the estimated position(s) which of the one or more people are within the spatial boundary, and generating the one or more video signals based only on the crop regions corresponding to the one or more people within the spatial boundary.

[0009] The method may further comprise the step of transmitting the or each video signal to a receiver, which may be a second video conferencing endpoint connected to the first video conferencing endpoint via a computer network.

[0010] The steps of the method may be performed in any order as appropriate. For example, the step of receiving data defining a spatial boundary may occur after the step of capturing an image of the field of view.

[0011] Framing may refer to extracting an area, e.g., a crop area, of the captured image that includes the person determined to be within the spatial boundary. The frame or crop area may be smaller than the originally captured image, and the framed person may be centered within the extracted area. In some examples, a crop area or one of the crop areas may include only a single person. In some examples, a crop area or one of the crop areas may include multiple people, each determined to be within the spatial boundary. In one example, a single crop area is extracted that includes all of the people determined to be within the spatial boundary.

[0012] The method may further include a verification mode in which each person in an image within the camera's field of view is labeled according to whether it is inside or outside a spatial boundary, and the labeled image is presented to a user for verification. The user can then modify the data defining the spatial boundary to ensure that all people framed are within the spatial boundary.

[0013] The step of estimating the position of the or each person may be performed by measuring the distance between one or more pairs of facial landmarks for each person. For example, the estimation may be performed by taking the average distance between one or more pairs of facial landmarks for a human, detecting these landmarks in the captured image, calculating the distance between the landmarks in the image, estimating the position of the person relative to the camera based on the camera imaging geometry and camera parameters, and estimating the distance from a plurality of distances calculated from each pair of facial landmark features.

[0014] The step of estimating the distance may include estimating an orientation of the person's face relative to the camera and selecting a pair of facial landmarks to be used to estimate the position based on the estimated orientation.

[0015] The step of estimating the position of the or each person may include estimating a camera orientation using one or more accelerometers in the videoconferencing endpoint.

[0016] The step of estimating the position of the or each person may include the use of one or more distance sensors within the videoconferencing endpoint.

[0017] The spatial boundary is defined at least in part as a distance from the location of the camera. The distance may be a radial distance that effectively forms a circular boundary at the floor. In another example, the spatial boundary defines how far from the side and how far forward from the camera to form a rectangular boundary at the floor. The spatial boundary may be defined at least in part by the angular range of the captured images.

[0018] The method may include a user input step in which the user provides data defining the spatial boundary. The user may provide the data via a user interface, for example by defining a distance to the side or forward from the camera via the user interface. The user may provide the data by causing the videoconferencing endpoint to enter a data entry mode in which the videoconferencing endpoint tracks the location of the user, and the user requests the videoconferencing endpoint to define the spatial boundary using the user's location or locations.

[0019] The method may be performed on a video stream, whereby the position of the or each person within the field of view of the camera is tracked and the step of generating one or more video signals is repeated for multiple images of the field of view.

[0020] In a second aspect, an embodiment of the present invention provides a video conferencing endpoint including a video camera configured to capture an image indicative of a field of view, and a processor, the processor comprising: receiving data defining a spatial boundary within a field of view, the spatial boundary being defined at least in part by a distance from the video camera; acquiring an image of the field of view from a video camera; identifying one or more persons within a field of view of a video camera; estimating the position of the or each person within a field of view of a video camera; generating one or more video signals including one or more crop regions corresponding to one or more people determined to be within the spatial boundary for transmission to a receiver; The present invention provides a video conferencing endpoint configured to:

[0021] The videoconferencing endpoint of the second aspect may be configured to perform any one of the features of the method described in the first aspect, or any combination thereof where suitable.

[0022] In a third aspect, an embodiment of the invention provides a computer-implemented method for estimating a distance from a person to a camera, the method comprising: (a) acquiring an image of a person with a camera; (b) identifying a human face region present on the camera; (c) measuring the distance between each of a plurality of pairs of human facial landmarks; (d) estimating a distance of the person from the camera using each of the measured distances; (e) identifying a maximum estimated distance and / or a minimum estimated distance in step (d); (f) estimating the position of the person relative to the camera based on the identified maximum and / or minimum distances; The present invention provides a computer-implemented method, comprising:

[0023] In a fourth aspect, an embodiment of the present invention provides a videoconferencing endpoint configured to perform the method of the third aspect.

[0024] The present invention includes combinations of the described aspects and optional features unless such combinations are clearly impermissible or explicitly avoided.

[0025] Further aspects of the invention provide a computer program comprising code which, when executed on a computer, causes a computer to perform the method of the first and / or third aspect, a computer readable medium storing a computer program comprising code which, when executed on a computer, causes a computer to perform the method of the first and / or third aspect, and a computer system programmed to perform the method of the first and / or third aspect. [Brief description of the drawings]

[0026] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which: [Figure 1] 1 illustrates a video conferencing endpoint. [Diagram 2] 1 shows a flow chart of a computer-implemented method. [Diagram 3] 2 illustrates a video conferencing suite including the video conferencing endpoints of FIG. 1. [Figure 4] 2 illustrates different video conferencing suites that include the video conferencing endpoints of FIG. 1; [Diagram 5] 4 shows a verification image displayed to the user. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0027] Aspects and embodiments of the present invention will now be discussed with reference to the accompanying drawings. Further aspects and embodiments will be apparent to those skilled in the art.

[0028] Figure 1 shows a videoconferencing endpoint 100. The endpoint includes a processor 2 connected to a volatile memory 4 and a non-volatile memory 6. Either or both of the volatile memory 4 and the non-volatile memory 6 store machine-executable instructions that, when executed in the processor, cause the processor to perform the method discussed with reference to Figure 2. The processor 2 is also connected to one or more video cameras 102; in this example there is a single camera, but there may be multiple cameras providing different fields of view or capture modes (e.g. frequency ranges). The processor is also connected to one or more microphones 12, and a human machine interface 14 (e.g. a keyboard or a touch-enabled display) that allows a user to input data. The processor is also connected to a network interface 8 that allows data to be transmitted over a network.

[0029] FIG. 2 shows a flow chart of the computer-implemented method. In a first step 202, the processor receives data defining a spatial boundary within the field of view of the camera or cameras 102. This data may be received, for example, via the human machine interface 14 or via the network interface 8. This data may, for example, identify the maximum distance from the camera (e.g., in meters) that the spatial boundary bounds. The data may also, for example, identify the maximum angle from the camera over which the spatial boundary extends. In one example, the data is received by the user having the videoconferencing endpoint enter a data entry mode in which the processor 2 tracks the user's location with the camera(s) 102. The user then requests the videoconferencing endpoint to define the vertices or boundaries of the spatial boundary using the user's current location. This request may, for example, be by the user making a gesture in a predefined manner (e.g., crossing their arms in an "X" shape). The user may then move to another location and repeat the gesture to define a second vertex or boundary, and so on.

[0030] After the processor receives the data, the method moves to step 204 where an image is taken by the camera of a field of view that includes spatial boundaries. The processor then identifies all people in the field of view in step 206. This identification of people may be done using, for example, a machine learning model trained to identify people in images. In some examples, a "you only look once" (or YOLO) object detection algorithm, or a computer vision Haar feature-based cascaded classifier, or a trained convolutional neural network such as a histogram of oriented gradients may be used to identify people in the image. The processor increments a counter j to indicate the number of people identified in the field of view of the camera. The processor then enters into a loop defined by steps 208-216. In step 208, the position of person i is estimated in the field of view of the camera.

[0031] Estimating the position or location of a person in the field of view is done in four steps in some examples: (i) estimating the distance from the person's face to the camera, (ii) calculating the direction to the person's face relative to the camera horizontal position, (iii) calculating the orientation of the camera by using one or more accelerometers at the endpoint, and (iv) calculating the direction of the person's face relative to the plane of the room floor in the field of view. Steps (i)-(iii) can be done in any order. The first step can be done by various methods including (a) using a time-of-flight sensor, (b) using stereo vision from two or more cameras, (c) using a trained machine learning algorithm on the image, (d) detecting a face in the image and using the size of the bounding box of the face, (e) detecting a face and then detecting facial landmarks such as eyes, nose, mouth, etc. and estimating the distance using a pre-trained machine learning model, and (f) detecting key features of a person such as head, ears, torso, etc. and using a pre-trained machine learning model that assumes a constant distance between at least some of the key features.

[0032] It is also possible to estimate the person's position using a variation of (e). The distances between pairs of facial landmarks vary within 10% across populations. Some examples of these distances include the distance between the eyes, the distance between one eye and the tip of the nose, the distance between one eye and the mouth, the distance between the top of the forehead and the chin, and the overall width of the face. In the captured image, these landmarks are projected onto the camera focal plane, and therefore the distances between the landmarks in the captured image depend on the camera's viewing angle of the face. If the person turns their face to the side with respect to the camera field of view, most of the above distances will be shorter. However, some will not be, including (for example) the face length or the distance from a visible eye to the mouth. Similarly, if the person looks upwards, the projected face length in the image will be shorter, but the face width and eye distance will remain the same. If the person rotates their face but keeps their face frontal to the camera, the distances between the landmarks, such as the eye distance, will remain the same. Assuming that the distance between landmarks is smaller than the distance from the face to the camera, the imaging of the camera makes it possible to derive equations relating the distance between two landmarks in the real world and their distance in the image in pixel units. These equations, sometimes called equivalence equations, exhibit properties of triangle ratios.

[0033] For example, let f be the focal length of the camera (in meters) and d real Let be the real-world distance between two facial landmarks, and d image Let be the distance in the image between two facial landmarks in pixel length units, pixelSize be the size of a pixel (in meters), and d be the distance from the person to the camera, then we can derive:

number

number

[0034] The position of a person's face relative to the camera position is uniquely identified by knowing the distance and direction. The direction can be described by angles such as pan and tilt. The direction of the face relative to the camera horizontal plane may be calculated from the pixel position of the face relative to the center of the image. For example, if cx is the position of the face in the image relative to the center pixel, the pan angle for a telephoto lens can be calculated as follows:

number

number

[0035] Videoconferencing endpoints are often mounted at an angle, either upwards or downwards, relative to the floor. The camera orientation may be calculated from an accelerometer in the endpoint, which senses gravity, to allow the tilt angle to be derived. The orientation relative to the floor is derived from the angles above. For example, the pan angle relative to the floor is equal to the pan angle relative to the camera horizontal, while the tilt angle relative to the floor is equal to the sum of the tilt angle relative to the camera horizontal and the tilt angle of the camera.

[0036] Once the person's position has been estimated, the method proceeds to step 210 where the processor determines whether person i is within a predefined spatial boundary. If so, i.e., "yes", the method proceeds to step 212 where the person is added to a framing list (i.e., a list containing the person or people to be framed in one or more crop regions). The method then proceeds to step 214 where the i counter is incremented. If the person is determined to be outside the spatial boundary, i.e., "no", the method proceeds directly to step 214 and step 212 is not performed.

[0037] Once the counter is incremented, the processor determines in step 216 whether i=j, i.e., whether the positions of all identified persons have been estimated and compared to the boundaries. If not, i.e., "no", the method returns to step 208 and continues looping. Note that in one example, the method may first loop through all persons identified in step 206 to estimate their positions, and then loop through each estimated position to determine if they are within the spatial boundaries. The method may then loop through all persons determined to be within the spatial boundaries and frame them. Once all persons have been estimated and the decision has been made as to whether to frame them, i.e., "yes", the method moves to step 218, where a crop region or each crop region is extracted that includes one or more of the persons in the framing list. These crop regions are then used to generate one or more single video streams, each video stream including its respective crop region, or a composite video stream including multiple crop regions. These are transmitted in step 220.

[0038] In an alternative method, all of the people identified in step 206 are first framed, i.e., a crop region is extracted for each person identified in step 206. Next, the method identifies each person within a spatial boundary and separates the crop regions that include the person within the spatial boundary from the remaining crop regions. Then, only the crop regions that include the person within the spatial boundary are used.

[0039] FIG. 3 shows a videoconferencing suite including the videoconferencing endpoint 100 of FIG. 1. The camera 102 captures a field of view 106 (indicated by a dashed line) that includes a first room 104 and a second room 110. The first and second rooms are separated by a glass wall 112, in this example the room 104 is the videoconferencing suite and the room 110 is an office. A spatial boundary 108 (indicated by a dotted line) is defined as the maximum distance from the camera. In this example, this means that people 114a-114d are within the spatial boundary, while person 116 (who is within the field of view 106 of the camera 102 but not in the first room 104) is not within the spatial boundary. Thus, people 114a-114d can be framed by the videoconferencing endpoint 100 and person 116 can be excluded.

[0040] Figure 4 illustrates a different videoconferencing suite including the videoconferencing endpoints of Figure 1. Like features are indicated with like reference numerals. In contrast to the example illustrated in Figure 3, here the spatial boundaries are not only defined by a maximum distance 108, but are further defined by a maximum angular range 408 of the image. By appropriately defining the maximum angular range, people 114a-114b can be defined as being within the spatial boundaries, while person 116 can be excluded from the spatial boundaries.

[0041] FIG. 5 shows a verification image displayed to the user. A graphical indication is provided and associated with each person that a person is inside or outside the spatial boundary. In this example, a check is placed near the person who is inside the spatial boundary, while a cross is placed near the person who is outside the spatial boundary. Other graphical indications may be provided, such as a bounding box around only the people found to be within the spatial boundary, or a bounding box around all detected people but with different colors for people inside and outside the boundary. This may allow the user to adjust the data defining the spatial boundary to appropriately exclude or include people to be framed.

[0042] The features disclosed in this description, or the following claims, or the accompanying drawings, which are expressed in a particular form or in terms of means for performing a disclosed function, or methods or processes for obtaining the disclosed results, may be utilized to realize the invention in its various forms, either separately or in any combination of such features, as appropriate.

[0043] While the present invention has been described in conjunction with the exemplary embodiments above, numerous equivalent modifications and variations will be apparent to those skilled in the art in light of this disclosure. Accordingly, the exemplary embodiments of the present invention set forth above are considered to be illustrative and not limiting. Various changes may be made to the described embodiments without departing from the spirit and scope of the present invention.

[0044] For the avoidance of any doubt, any theoretical explanations provided herein are provided for the purpose of enhancing the understanding of the reader, and the inventors do not wish to be bound by any of these theoretical explanations.

[0045] Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0046] Throughout this specification, including the claims which follow, unless the context requires otherwise, the words "comprise" and "include", as well as variations such as "comprises", "comprising" and "including", are understood to imply the inclusion of a stated integer or step or group of steps, but not the exclusion of any other integer or step or group of integers or steps.

[0047] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values ​​are expressed as approximations, by use of "about," it will be understood that the particular value forms another embodiment. The term "about" with respect to numerical values ​​is optional and may mean, for example, + / - 10%.

Claims

1. 1. A computer-implemented method of operating a videoconferencing endpoint, the videoconferencing endpoint including a video camera that captures images indicative of a field of view, the method comprising: receiving data defining a spatial boundary within the field of view, the spatial boundary being defined at least in part by a distance from the video camera; taking an image of the field of view; identifying one or more people within the field of view of the video camera; estimating a position of the or each person within the field of view of the video camera; generating one or more video signals for transmission to a receiver, the video signals including one or more crop regions corresponding to one or more people determined to be within the spatial boundary; 4. A computer-implemented method comprising:

2. generating the one or more video signals determining from the estimated location(s) that at least one of the one or more persons is within the spatial boundary; framing the one or more people determined to be within the spatial boundary to generate respective crop regions; The computer-implemented method of claim 1 , comprising:

3. A computer-implemented method as claimed in claim 1 or 2, comprising the step of transmitting the or each video signal to said receiver.

4. labeling each person in the image within the field of view of the video camera according to whether they are inside or outside the spatial boundary; Present the labeled image to a user for verification. The computer-implemented method of claim 1 further comprising a validation mode.

5. 2. The computer-implemented method of claim 1, wherein estimating the position of the or each person is performed by measuring the distance between one or more pairs of facial landmarks for each of the persons.

6. 6. The computer-implemented method of claim 5, wherein distances between multiple pairs of facial landmark features are measured, each distance is used to estimate a distance of the person from the video camera, and a maximum and / or minimum of the estimated distances is used to estimate the position of the or each person.

7. 7. The computer-implemented method of claim 5 or 6, wherein the step of estimating the distance comprises estimating an orientation of the person's face relative to the camera and selecting a pair of facial landmarks to be used to estimate the position based on the estimated orientation.

8. 2. The computer-implemented method of claim 1, wherein estimating the position of the or each person comprises estimating an orientation of the camera using one or more accelerometers in the videoconferencing endpoint.

9. The computer-implemented method of claim 1 , wherein estimating the location of the or each person comprises use of one or more distance sensors within the videoconferencing endpoint.

10. The computer-implemented method of claim 1 , wherein the spatial boundaries are further defined, at least in part, by an angular range of the captured images.

11. The computer-implemented method of claim 1 , wherein the method includes a user input step in which a user provides the data defining the spatial boundary.

12. The computer-implemented method of claim 11 , wherein the user provides the data via a user interface.

13. 12. The computer-implemented method of claim 11, wherein the user prepares the data by causing the videoconferencing endpoint to enter a data entry mode in which the videoconferencing endpoint tracks the user's location, and the user requests the videoconferencing endpoint to define the spatial boundary using one or more locations of the user.

14. 1. A video conferencing endpoint including a video camera configured to capture an image indicative of a field of view, and a processor, the processor comprising: receiving data defining a spatial boundary within the field of view, the spatial boundary being defined at least in part by a distance from the video camera; acquiring an image of the field of view from the video camera; identifying one or more people within the field of view of the video camera; estimating a position of the or each person within the field of view of the video camera; generating one or more video signals including one or more crop regions corresponding to one or more people determined to be within the spatial boundary for transmission to a receiver; A video conferencing endpoint that is configured to

15. generating the one or more video signals determining from the estimated location(s) that at least one of the one or more persons is within the spatial boundary; framing the one or more people determined to be within the spatial boundary to generate respective crop regions; 15. The video conferencing endpoint of claim 14, comprising:

16. 16. The video conferencing endpoint of claim 14 or 15, wherein the video conferencing endpoint is connected to a receiver via a network, and the processor is configured to transmit the one or more video signals to the receiver.

17. The processor, labeling each person in the image within the field of view of the camera according to whether they are inside or outside the spatial boundary; Present the labeled image to a user for verification.

15. The videoconferencing endpoint of claim 14, configured to perform a verification mode.

18. 15. The video conferencing endpoint of claim 14, wherein the processor is configured to estimate the position of the or each person by measuring a distance between one or more pairs of facial landmark features of the or each person.

19. 20. The video conferencing endpoint of claim 18, wherein the processor is configured to measure a plurality of distances between a plurality of pairs of facial landmark features, estimate a distance of the person from the video camera using each of the measured distances, and estimate the position of the or each person using a maximum and / or minimum estimated distance among the estimated distances.

20. 1. A computer-implemented method for estimating a distance from a person to a camera, the method comprising: (a) capturing an image of the person with the camera; (b) identifying a facial region of the person present in the image; (c) measuring the distance between each of a plurality of pairs of the human facial landmarks; (d) estimating a distance of the person from the camera using each of the measured distances; (e) identifying a maximum estimated distance and / or a minimum estimated distance in step (d); (f) estimating a position of the person relative to the camera based on the identified maximum and / or minimum distances; 4. A computer-implemented method comprising: