Detecting and framing objects of interest in a teleconference

By detecting and estimating the head posture and gaze of video conference participants using optical sensors, and identifying regions of interest, the problem of inaccurate view selection in existing technologies is solved, resulting in more efficient view display and interactive effects.

CN114846787BActive Publication Date: 2026-07-31HEWLETT PACKARD DEVELOPMENT COMPANY LP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEWLETT PACKARD DEVELOPMENT COMPANY LP
Filing Date
2021-01-26
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing video conferencing systems struggle to accurately capture the view of meeting participants when selecting and framing the best view, resulting in poor interactive experiences.

Method used

By receiving image data frames through optical sensors, the system detects meeting participants, estimates their head posture and gaze direction, identifies regions of interest, determines that objects within overlapping areas are objects of interest, and optimizes view selection and presentation.

Benefits of technology

It improves the accuracy of view selection and interactive effects in video conferencing, enhances the view display quality of remote endpoints, reduces distractions, and optimizes meeting efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114846787B_ABST
    Figure CN114846787B_ABST
Patent Text Reader

Abstract

A method for view selection in a teleconference environment includes: receiving image data frames from an optical sensor such as a camera; detecting one or more conference participants within the image data frames; and identifying regions of interest (ROIs) for each of the conference participants. Identifying ROIs involves: estimating the participants' head posture to determine where most participants are looking and determining if an object exists in that area. If a suitable object is present in the area being viewed by a participant, such as a whiteboard or another person, image data corresponding to that object is displayed on a display device or sent to a remote teleconference endpoint.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to video conferencing and specifically to accurately estimating the head posture of meeting participants. Background Technology

[0002] During a video conference, people at one video conference endpoint interact with people at one or more other video conference endpoints. Systems exist that capture views of conference participants from different angles. Attempts to create systems that automatically select and frame the best view for transmission to remote endpoints, primarily based on who is currently speaking, are not entirely satisfactory. Therefore, there is room for improvement in this field. Summary of the Invention

[0003] According to examples of this disclosure, a method for view selection in a teleconference environment includes: receiving image data frames from an optical sensor such as a camera, detecting one or more conference participants within the image data frames, and identifying regions of interest (ROIs) for each of the conference participants. Identifying ROIs includes: estimating the participants' head postures to determine where the participants are looking and determining whether an object exists in that area. If a suitable object is in the area that a participant is looking at, such as a whiteboard or someone else, image data corresponding to that object will be displayed on a display device or sent to a remote teleconference endpoint, or presented in some other way.

[0004] One example of this disclosure is a method for view selection in a teleconference environment, comprising: receiving an image data frame from an optical sensor; detecting one or more conference participants within the image data frame; identifying regions of interest for each of the one or more conference participants, wherein identifying regions of interest for each of the one or more conference participants includes: estimating a head pose from a first participant among the one or more conference participants; determining that most regions of interest overlap in an overlapping area; detecting objects within the overlapping area; determining that the objects within the overlapping area are objects of interest; and presenting a view containing the objects of interest.

[0005] Another example of this disclosure includes a teleconference endpoint, comprising:

[0006] An optical sensor is configured to receive image data frames; a processor is coupled to the optical sensor, wherein the processor is configured to: detect one or more conference participants within the image data frames; identify regions of interest for each of the one or more conference participants by estimating the head pose of a first participant from the one or more conference participants; determine that most of the regions of interest overlap in an overlapping region; detect objects within the overlapping region; determine that the objects within the overlapping region are objects of interest; and present a view containing the objects of interest.

[0007] Another example of this disclosure includes a non-transitory computer-readable medium storing instructions executable by a processor, the instructions including instructions for performing the following operations: receiving an image data frame from an optical sensor; detecting one or more conference participants within the image data frame; identifying regions of interest for each of the one or more conference participants, wherein identifying regions of interest for each of the one or more conference participants includes: estimating the head pose of a first participant among the one or more conference participants; determining that more regions of interest overlap in an overlapping region; detecting objects within the overlapping region; determining that the objects within the overlapping region are objects of interest; and presenting a view containing the objects of interest in a transmission to a remote endpoint. Attached Figure Description

[0008] For illustration, certain examples described in this disclosure are shown in the accompanying drawings. In the drawings, the same numerals indicate the same elements throughout. The full scope of the invention disclosed herein is not limited to the precise arrangements, dimensions, and apparatus shown. In the drawings:

[0009] Figure 1 The illustration shows a video conferencing endpoint according to an example of this disclosure;

[0010] Figure 2A The diagram shows Figure 1 The aspects of video conferencing endpoints;

[0011] Figure 2B An aspect of a camera according to an example of this disclosure is illustrated;

[0012] Figures 3A-3E The illustration shows an example of receiving and evaluating an image data frame according to this disclosure;

[0013] Figure 4A The illustration shows a method for determining an object of interest according to an example of this disclosure;

[0014] Figure 4B The illustration shows another method for determining an object of interest according to an example of this disclosure.

[0015] Figure 5 The illustration shows a focus estimation model based on an example of this disclosure;

[0016] Figure 6 The illustration shows a method for selecting objects of interest according to an example of this disclosure;

[0017] Figures 7A-7F The diagram illustrates the following: Figure 6 The method selects aspects of the object of interest;

[0018] Figures 8A-8B The diagram illustrates the following: Figure 6 The method selects other aspects of the object of interest; and

[0019] Figure 9 An electronic device that can be used to practice the concepts and methods of this disclosure is illustrated. Detailed Implementation

[0020] In the accompanying drawings and description thereof, certain terms are used for convenience only and should not be considered as examples limiting this disclosure. In the drawings and the following description, the same numerals consistently denote the same elements.

[0021] the term

[0022] Throughout this disclosure, terminology is used in a manner consistent with that used by those skilled in the art, for example:

[0023] The centroid or geometric center of a planar graph is the arithmetic mean position of all points in the graph.

[0024] A normal is an object perpendicular to a given object, such as a line or vector. In two dimensions, the normal to a curve at a given point is a line perpendicular to the tangent to the curve at that point. In three dimensions, the normal to a face at a point is a vector perpendicular to the plane tangent to that face at that point.

[0025] discuss

[0026] In one or more examples of this disclosure, the object of interest is determined based on multiple factors. In at least one example of this disclosure, the video conferencing device can detect and focus on an active speaker. One or more microphone arrays can be used to determine the direction from the video conferencing device to the active speaker. In one or more examples of this disclosure, one or more cameras are used to locate the face of the active speaker. In some examples, sound source localization is used to detect the active speaker. In some examples, body detection is used to detect the active speaker. In some examples, lip movement detection is used to locate the current speaker. In at least one example, the current speaker is located, and one or more cameras can automatically point at him or her. A view of the active speaker can be captured for transmission to another endpoint, and the active speaker can be tracked during a video conference.

[0027] In some examples of this disclosure, additional bases are utilized for selecting one or more views (or portions of views) for presentation. In at least one example, the chart at the endpoint becomes an object of interest when the speaker refers to it. In at least one example, the assembly participant and speaker at the endpoint become objects of interest when the speaker addresses the participants. In at least one example, the object becomes an object of interest when the speaker makes a gesture pointing to it. In at least one example, the assembly participant and speaker at the endpoint become objects of interest when the speaker discusses the assembly participant in the third person. According to examples of this disclosure, one or more views depicting the objects of interest are transmitted to a remote endpoint for viewing.

[0028] This disclosure relates to optimizing how to box an object of interest. At least one example of this disclosure relates to determining where to position the object of interest within the box. In at least one example, when the object of interest is a person with at least one eye in the field of view of the capturing camera, the degree to which the person is placed away from the centroid of the presented box is a function of the degree to which the person's gaze is deviated from the capturing camera.

[0029] In at least one example of this disclosure, an object or person becomes an object of interest when a majority of participants at an endpoint are viewing it. In at least one example of this disclosure, an object or person becomes an object of interest when multiple participants at an endpoint are viewing it.

[0030] In at least one example of this disclosure, head pose estimation is used as a clue to find the object or person the participant is looking at. In at least one example, eye gaze estimation is used as a clue to find the object or person the participant is looking at. In at least one example of this disclosure, head pose estimation and eye gaze estimation are used as clues to find the object or person the participant is looking at. In at least one example, the voting module acquires head pose and eye gaze estimation data and searches for “hotspot areas” that currently attract attention. In some examples, the object detection module determines whether an object exists in the “hotspot area.” The object can be a person or object, such as a whiteboard, screen, activity chart, or product.

[0031] In at least one example of this disclosure, a decision will be made to display a view containing the object of interest. Displaying this view may include switching from an earlier view. Switching views may include switching between cameras, panning or zooming (mechanically or electronically) one of the cameras, switching to a content stream, switching to the smart board's output, and switching to a dedicated whiteboard camera.

[0032] In at least one example of this disclosure, the focus estimation model is used to determine where a person is looking within a single frame or a series of frames. In this example, focus estimation is performed by a trained neural network to take an input image and output a focus map. The focus map is a probability distribution that indicates how likely a person at a particular location is to be interested in a particular area.

[0033] The technical benefits of identifying areas of interest within a meeting space include helping to determine what type of meeting space makes a meeting more effective, how to reduce distractions, and how long to schedule a meeting.

[0034] According to the examples in this disclosure, once the object of interest is identified, a determination is made regarding how to display the object of interest in an optimized manner.

[0035] Figure 1 A video conferencing endpoint 100 according to an example of this disclosure is illustrated. The video conferencing device or endpoint 100 communicates with one or more remote endpoints 102 via a network 104. Components of endpoint 100 include an audio module 106 with an audio codec 108 and a video module 110 with a video codec 112. Modules 106 and 110 are operatively coupled to a control module 114 and a network module 116. In one or more examples, endpoint 100 includes exactly one wide-angle electronic pan-tilt-zoom camera. In some examples, when a view object is zoomed in, a sub-part of the captured image containing the object is rendered, while other parts of the image are not rendered.

[0036] During a video conference, two or more cameras (e.g., cameras 118 and 120) capture video and provide the captured video to video module 110 and codec 112 for processing. In at least one example of this disclosure, one camera (e.g., 118) is a smart camera and one camera (e.g., 120) is not a smart camera. In some examples, two or more cameras (e.g., cameras 118 and 120) are cascaded such that one camera controls some or all of the operations of another camera. In some examples, two or more cameras (e.g., cameras 118 and 120) are cascaded such that data captured by one camera (e.g., by control module 114) is used to control some or all of the operations of another camera. Furthermore, one or more microphones 122 capture audio and provide the audio to audio module 106 and codec 108 for processing. These microphones 122 may be desktop or ceiling microphones, or they may be part of a microphone box, etc. In one or more examples, microphones 122 are closely coupled to one or more cameras (e.g., cameras 118 and 120). The audio captured by these microphones 122 at endpoint 100 is primarily used for conference audio.

[0037] Endpoint 100 also includes a microphone array 124, wherein subarrays 126 and 128 are arranged orthogonally. Microphone array 124 also captures audio and provides it to audio module 22 for processing. In some examples, microphone array 124 includes vertically and horizontally arranged microphones for determining the location of an audio source (e.g., a person speaking). In some examples, endpoint 100 uses the audio from array 124 primarily for camera tracking purposes and not for conference audio. In some examples, endpoint 100 uses the audio from array 124 for both camera tracking and conference audio.

[0038] After capturing audio and video, endpoint 100 encodes the audio and video according to encoding standards such as MPEG-1, MPEG-2, MPEG-4, H.261, H.263, and H.264. Then, network module 116 outputs the encoded audio and video to remote endpoint 102 via network 104 using an appropriate protocol. Similarly, network module 116 receives conference audio and video from remote endpoint 102 via network 104 and transmits the received audio and video to their respective codecs 108 / 112 for processing. Endpoint 100 also includes a speaker 130 for outputting conference audio and a display 132 for outputting conference video.

[0039] In at least one example of this disclosure, endpoint 100 uses two or more cameras 118, 120 in an automatic and coordinated manner to dynamically process video and views of the video conferencing environment. In some examples, the first camera (e.g., 118) is a fixed or room-view camera, and the second camera 120 is a controlled or person-view camera. Using the room-view camera (e.g., 118), endpoint 100 captures video of the room or at least typically a wide or zoomed-out view of the room including all video conferencing participants 121 and some of their surroundings.

[0040] According to some examples, endpoint 100 uses a human-view camera (e.g., 120) to capture video of one or more participants, including one or more current speakers, in a compact or magnified view. In at least one example, the human-view camera (e.g., 120) can pan, tilt, and / or zoom.

[0041] In one arrangement, the human-view camera (e.g., 120) is a steerable pan-tilt-zoom (PTZ) camera, while the room-view camera (e.g., 118) is an electronic pan-tilt-zoom (EPTZ) camera. Therefore, the human-view camera (e.g., 120) can be steered, while the room-view camera (e.g., 118) cannot. In at least one example, both camera 118 and camera 120 are EPTZ cameras. In at least one example, camera 118 is associated with sound source locator module 134. In fact, both cameras 118 and 120 can be steerable PTZ cameras.

[0042] In some examples, endpoint 100 will alternate between a tight view of the speaker and a wide view of the room. In some examples, endpoint 100 will alternate between two different tight views of the same or different speakers. In some examples, endpoint 100 will capture a first view of a person with one camera and a second view of the same person with another camera, and determine which view is better for sharing with remote endpoint 102.

[0043] In at least one example of this disclosure, endpoint 100 outputs video from only one of the two cameras 118, 120 at any given time. As the video conference progresses, the output video from endpoint 100 can switch from the view of one camera to the view of the other camera. According to some examples, when no participant is speaking in the person view while one or more participants 121 are speaking, system 100 outputs a room view.

[0044] According to the example, endpoint 100 can transmit video from two cameras 118 and 120 simultaneously, and endpoint 100 can allow remote endpoint 102 to decide which view to display, or determine how one view will be displayed relative to another view in a specific way. For example, one view can be composited into a picture-in-picture of another view.

[0045] In one or more examples, endpoint 100 uses an audio-based locator 134 and a video-based locator 136 to determine the location of participant 121 and a frame view of the environment and participant 121. Control module 114 uses the audio and / or video information from these locators 134, 136 to crop one or more captured views such that one or more sub-portions of the captured views are displayed on display 132 and / or transmitted to remote endpoint 102. In some examples, commands to one or both cameras 118, 120 are implemented by an actuator or local control unit 137 having motors, servos, etc., to mechanically manipulate one or both cameras 118, 120. In some examples, such camera commands may be implemented as electronic signals by one or both cameras 118, 120.

[0046] In some examples, to determine which camera view to use and how to configure the view, control module 114 uses audio information obtained from audio-based locator 134 and / or video information obtained from video-based locator 136. For example, control module 114 uses audio information from horizontally and vertically arranged microphone subarrays 126, 128 processed by audio-based locator 134. Audio-based locator 134 uses speech detector 138 to detect speech in the captured audio from subarrays 126, 128 to determine the current location of the participant. Control module 114 uses the determined location to maneuver the person-view camera toward that location. In some examples, control module 114 uses video information captured using cameras 118, 120 and processed by video-based locator 136 to determine the location of participant 121, determine a bounding box for the view, and maneuver one or more of the cameras (e.g., 118, 120). In other examples, none of the cameras are physically maneuverable.

[0047] A wide view from one camera (e.g., 118) can provide a background for a zoom view from another camera (e.g., 120), such that while the video from the other camera (e.g., 120) is being adjusted, the participant 121 at remote endpoint 102 sees the video from one camera (e.g., 118). In some examples, the transition between the two views from cameras 118, 120 can be softened and blended to avoid abrupt cuts when switching between camera views. In some examples, a switch from the first view to the second view for transmission to remote endpoint 102 will not occur until the active participant 121 has been shown in the second view for a minimum amount of time. In at least one example of this disclosure, the minimum amount of time is one second. In at least one example, the minimum amount of time is two seconds. In at least one example, the minimum amount of time is three seconds. In at least one example, the minimum amount of time is four seconds. In at least one example, the minimum amount of time is five seconds. In other examples, other minimum values ​​(e.g., 0.5-7.0 seconds) are used, depending on factors such as the size of the meeting room, the number of participants 121 at endpoint 100, the cultural details of participants 140 at remote endpoint 102, and the size of one or more monitors 132 displaying the captured view.

[0048] Figure 2 illustrates aspects of a video conferencing endpoint 200 (e.g., 100) according to an example of the present disclosure. Endpoint 200 includes a speaker 130, a camera 202 (e.g., 118, 120), and a microphone 204 (e.g., 122, 124). Endpoint 200 also includes a processing unit 206, a network interface 208, a memory 210, and an input / output interface 212, all connected by a bus 101.

[0049] Memory 104 can be any conventional memory (such as SDRAM) and can store modules 216 in the form of software and firmware for controlling endpoint 200. In addition to audio and video codecs (108, 112) and other modules discussed previously, module 216 may include an operating system, a graphical user interface (GUI) enabling users to control endpoint 200, and algorithms for processing audio / video signals and controlling camera 202. In at least one example of this disclosure, one or more of the cameras 202 may be panoramic cameras.

[0050] Network interface 208 enables communication between endpoint 200 and remote endpoint (102). In one or more examples, interface 212 provides data transfer with local devices such as keyboards, mice, printers, overhead projectors, monitors, external speakers, additional cameras, and microphone boxes.

[0051] Camera 202 and microphone 204 capture video and audio respectively in a video conferencing environment and generate video and audio signals that are transmitted to processing unit 206 via bus 214. In at least one example of this disclosure, processing unit 206 processes the video and audio using algorithms in module 216. For example, endpoint 200 processes the audio captured by microphone 204 and the video captured by camera 202 to determine the position of participant 121 and to control and select views from camera 202. The processed audio and video can be sent to remote devices connected to network interface 208 and devices connected to general interface 212.

[0052] Figure 2B An aspect of a camera 202 according to an example of this disclosure is illustrated. The camera 202 has a lens 218. The lens 218 has a central region or centroid 220 and a focal length 222 between the center 352 of the lens 218 and the focal point 224 of the lens 218. The focal length 222 extends along the focal line 307 of the lens, which is orthogonal (perpendicular) to the lens 218.

[0053] Figures 3A-3E The illustration shows an example of receiving and evaluating image data frames according to this disclosure.

[0054] Figure 3A An image data frame 300 according to an example of this disclosure is illustrated. Frame 300 contains a view of a meeting room with multiple meeting participants 121.

[0055] Figure 3B The illustration shows the direction 302 that participant 121 is looking in. In at least one example, this assessment is based on estimating participant 121's head posture. In at least one example, this assessment is based on estimating participant 121's eye gaze.

[0056] Figure 3C The diagram illustrates this, based on... Figure 3B Based on the directional information obtained, some participants 121 are watching the first "hotspot area" 304 and some participants 121 are watching the second "hotspot area" 306.

[0057] Figure 3D The illustration shows how, once hotspot regions 304 and 306 are identified, a determination is made as to whether hotspot regions 304 and 306 contain objects. Figure 3D As can be seen, hotspot area 304 contains the first gathering participant and hotspot area 306 contains the second gathering participant. It is worth noting that while determining whether any participant 121 is currently speaking can be used when assessing who (or what) is currently the focus of interest, the example disclosed herein does not require determining who is the active speaker.

[0058] Figure 3E The illustration shows that once the hotspot area has been identified as corresponding to an object, a final determination is made regarding which object (person) is the object of interest 312. The object of interest 312 can be defined within a bounded area 314 of frame 300. Image data within the bounded area 314 can be presented, such as by transmitting image data to other participants 140 at a remote endpoint 102.

[0059] Figure 4A The illustration shows a method 401 for determining an object of interest according to an example of this disclosure. At step 402, an input frame (e.g., 300) is received, such as from camera 202. At step 404, head pose estimation and eye gaze estimation are used as clues to find objects or people the participant is looking at. At step 406, a voting module then acquires the estimated data and searches for “hotspot areas” that attract attention. Subsequently, an object detection module determines 408 whether an object exists in or near the “hotspot area.” As noted, the object can be a person (such as predicted by face detection operations), a whiteboard, a screen, a moving picture, a poster, etc. Subsequently, at step 410, a final decision is made (alone or in conjunction with other information), and a view containing the object of interest 312 is presented. Method 400 can end or return to step 402, in which another image data frame is received.

[0060] Figure 4BAn alternative method 401 for finding objects of interest according to an example of this disclosure is illustrated. At step 402, an input frame (e.g., 300) is received from camera 202. At step 412, a focus estimation model is used to assess where the participants' attention is focused. A trained neural network is used to perform focus estimation 412 to acquire the input image (e.g., 300) and output a focus map (not shown). The focus map contains a probability distribution indicating how likely a person (e.g., 121) at endpoint 100 is focusing their attention on a given area. After step 412 is completed, an object detection module determines 408 whether an object exists near the “hotspot area.” As noted, the object can be a person (e.g., predicted via face detection operations), a whiteboard, a screen, a moving picture, a poster, etc. Subsequently, in step 410, the object of interest is finally determined (alone or in conjunction with other information), and a view containing the object of interest 312 is presented. Method 400 may end or return to step 402, in which another image data frame is received.

[0061] Figure 5 The illustration shows a focus estimation model 500 according to an example of this disclosure. (As per...) Figure 4B The description describes a view frame 504 (e.g., 300) captured by camera 202. Image data 502 corresponding to the view frame is passed to a first convolutional layer 504 and a first modified linear activation function is applied. The modified output of the first convolutional layer is then passed to a first pooling layer 506. The output of the first pooling layer 506 is then passed to a second convolutional layer 508 and a second modified linear activation function is applied. The modified output of the second convolutional layer is then passed to a second pooling layer 510. The output of the second pooling layer 510 is then passed to a third convolutional layer 512 and a third modified linear activation function is applied. The modified output of the third convolutional layer 512 is then passed to a third pooling layer 514. The output of the third pooling layer 514 is then passed to a fourth convolutional layer 516 and a fourth modified linear activation function is applied. The modified output of the fourth convolutional layer 516 is then passed to a fourth pooling layer 518. The output of the fourth pooling layer 518 is then passed to a fifth convolutional layer 520 and a fifth modified linear activation function is applied. The corrected output of the fifth convolutional layer 520 includes a focus map 522. The focus map 522 is used to identify objects of interest in the manner discussed above (e.g., 312).

[0062] Figure 6A method 600 for selecting an object of interest 312 is illustrated. The method 600 begins by identifying (locating) 602 the object of interest 312 within an image data frame, for example, through method 400 and / or method 401. The object of interest 312 is initially selected within a default box (bounded area) 314. A determination 604 is then made regarding whether the object of interest 312 is a person. If the object of interest 312 is a person, the method 600 proceeds to estimate 606 the orientation of the person's head (or face). A portion of the original image data frame containing the object of interest 312 is then selected 608 for presentation, such as by sending the image data to a remote endpoint 104. According to the method 600, this portion of the frame is selected 608 to place the object of interest 312 in a manner that would be pleasing to participants in a shared view on a viewing display device (e.g., 132). On the other hand, if it is determined that the object of interest 312 is not a person, a default selection is used, in which the object of interest is substantially centered in the view.

[0063] Figures 7A-7F The illustration shows the positioning of object of interest 312 within the presented frame. Figure 7A In the frame, object of interest 312 is looking to the left, so it is correctly placed on the right side of the center. Figure 7B Object 312, which shares the same interest, is centered within the box, while Figure 7C The object of interest is located to the left of the center. Object of interest 312 is... Figure 7A The positioning within the box is the most visually pleasing of the three views in the top row of the page.

[0064] exist Figure 7D In the image, object of interest 312 is looking to the right of the frame and is placed in the center right, which is unpleasant. Figure 7E The same object of interest 312 is centered in the box, which is for... Figure 7E Improvements. Figure 7F The object of interest is correctly located to the left of the center and is furthest from the right side 702. Object of interest 312 is... Figure 7F The positioning within the box is the most visually pleasing of the three views in the middle row of the page.

[0065] Figures 8A-8B The illustration shows the positioning of object of interest 312 within the presented frame. Figure 8A In the diagram, object 312 is looking towards the right side 702 of the frame. The centroid 800 of object 312 (as illustrated) is located to the left of the center 802 of the frame, and is therefore appropriately positioned for viewing. On the other hand, in Figure 8BIn the diagram, object 312 is looking towards the right side 700 of the frame. The centroid 804 of object 312 (as illustrated) is located to the right of the center 802 of the frame and is therefore positioned appropriately for viewing.

[0066] Figure 9 An electronic device 900 (e.g., 100, 200) is illustrated that can be used to practice the described concepts and methods. The disclosed and described components can be integrated, in whole or in part, into a tablet computer, personal computer, mobile phone, and other devices using one or more microphones. As shown, device 900 may include a processing unit (CPU or processor) 920 and a system bus 910. System bus 910 interconnects various system components—including system memory 930, such as read-only memory (ROM) 940 and random access memory (RAM) 950—to processor 320. The processor may include one or more digital signal processors. Device 900 may include a cache 922 of high-speed memory that is directly connected to, close to, or integrated into processor 920. Device 900 copies data from memory 930 and / or storage device 960 to cache 922 for fast access by processor 920. In this way, the cache provides a performance improvement, avoiding latency for processor 920 while waiting for data. These and other modules can control or be configured to control processor 920 to perform various actions. Other system memory 930 may also be available for use. Memory 930 may include various different types of memory with different performance characteristics. Processor 920 may include any general-purpose processor and hardware or software modules (such as modules 1 (962), 2 (964), and 3 (966) stored in storage device 960) configured to control processor 920, as well as dedicated processors, where software instructions are incorporated into the actual processor design. Processor 920 may essentially be a completely independent computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0067] System bus 910 can be any of several types of bus architectures, including memory bus or memory controller, peripheral bus, and local bus using any of the various bus architectures. A basic input / output system (BIOS) stored in ROM 940, etc., can provide basic routines that facilitate the transfer of information between components within device 900, such as during startup. Device 900 also includes storage device 960, such as a hard disk drive, disk drive, optical disk drive, tape drive, etc. Storage device 960 can include software modules 962, 964, 966 for controlling processor 920. Other hardware or software modules are envisioned. Storage device 960 is connected to system bus 910 via a driver interface. Drivers and associated computer-readable storage media provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for device 900. In at least one example, the hardware module performing the function includes software components stored in a non-transitory computer-readable medium coupled to the hardware components required to perform the function—such as processor 920, bus 910, output device 970, etc.

[0068] To explain clearly, Figure 9 The device is presented as comprising individual functional blocks, including a functional block labeled "processor". The functionality represented by these blocks can be provided using shared or dedicated hardware, including but not limited to hardware capable of executing software, and hardware specifically built to operate equivalently to software executing on a general-purpose processor, such as processor 920. For example, Figure 9 The functionality of one or more processors illustrated herein may be provided by a single shared processor or multiple processors. (The term "processor" should not be construed as referring only to hardware capable of executing software.) One or more examples of this disclosure include microprocessor hardware and / or digital signal processor (DSP) hardware, read-only memory (ROM) 940 for storing software performing the operations discussed in one or more of the following examples, and random access memory (RAM) 950 for storing the results. Very large-scale integration (VLSI) hardware embodiments may also be used, as well as custom VLSI circuitry systems combined with general-purpose DSP circuitry (933, 935).

[0069] Examples of this disclosure also include:

[0070] 1. A method for view selection in a teleconference environment, comprising: receiving an image data frame from an optical sensor; detecting one or more conference participants within the image data frame; identifying a region of interest (ROI) for each of the one or more conference participants, wherein identifying the ROI for each of the one or more conference participants includes: estimating a head pose from a first participant among the one or more conference participants; determining that a majority of the ROIs overlap in an overlapping region; detecting an object within the overlapping region; determining that the object within the overlapping region is an object of interest; and presenting a view containing the object of interest.

[0071] 2. The method according to Example 1, wherein identifying the region of interest for each of the one or more meeting participants further includes estimating the gaze of a second participant among the one or more meeting participants.

[0072] 3. The method according to Example 2, wherein the first participant and the second participant are different.

[0073] 4. The method according to Example 1, wherein identifying the region of interest for each of the one or more meeting participants further includes: generating a focus map using a neural network.

[0074] 5. According to the method described in Example 1, determining that the object within the overlapping area is the object of interest further includes: determining that the object corresponds to a person.

[0075] 6. According to the method described in Example 5, determining that the object corresponds to a person includes: determining that the object corresponds to a person who has not spoken.

[0076] 7. The method according to Example 5, wherein presenting the view containing the object of interest comprises: determining the centroid corresponding to the object of interest; determining the gaze of the object of interest relative to a lens of an optical sensor for capturing the view containing the object of interest, the lens having a central region; determining that the gaze of the object of interest deviates from the normal of the central region by at least fifteen degrees; and positioning the object of interest within the view such that the centroid of the object of interest deviates from the centroid of the view.

[0077] 8. The method according to Example 7, wherein positioning the object of interest within the view such that the centroid of the object of interest is offset from the centroid of the view comprises: defining the object of interest within a rectangular bounded area having a horizontal width, and placing the object of interest within the rectangular bounded area such that the centroid of the object of interest is horizontally shifted from the boundary of the rectangular bounded area, the more pointed to by the gaze, by a distance corresponding to between one-half and two-thirds of the horizontal width. Other distances and ranges, such as between one-half and three-quarters, and between three-fifths and two-thirds, are covered within this disclosure.

[0078] 9. A teleconference endpoint, comprising: an optical sensor configured to receive an image data frame; and a processor coupled to the optical sensor, wherein the processor is configured to: detect one or more conference participants within the image data frame; identify a region of interest for each of the one or more conference participants by estimating a head pose from a first participant among the one or more conference participants; determine that a majority of the regions of interest overlap in an overlapping region; detect an object within the overlapping region; determine that the object within the overlapping region is an object of interest; and present a view containing the object of interest.

[0079] 10. The teleconference endpoint according to Example 9, wherein the processor is further configured to: identify the region of interest for each of the one or more conference participants by estimating the gaze from the one or a second participant among the conference participants.

[0080] 11. The teleconference endpoint according to Example 10, wherein the first participant and the second participant are different.

[0081] 12. The teleconference endpoint according to Example 9, wherein the processor is further configured to identify the region of interest for each of the one or more conference participants based on a focus map generated using a neural network.

[0082] 13. A teleconference endpoint according to Example 9, wherein the processor is further configured to determine that the object of interest corresponds to a person.

[0083] 14. A teleconference endpoint as described in Example 13, wherein the person is not an active speaker.

[0084] 15. A teleconference endpoint according to Example 13, wherein the processor is further configured to present the view containing the object of interest by: determining the centroid corresponding to the object of interest; determining the gaze of the object of interest relative to a lens of an optical sensor for capturing the view containing the object of interest, the lens having a central region; determining that the gaze of the object of interest deviates from the normal of the central region by at least fifteen degrees; and positioning the object of interest within the view such that the centroid of the object of interest deviates from the centroid of the view.

[0085] 16. A teleconference endpoint according to Example 15, wherein the processor is further configured to: position the object of interest within the view such that the centroid of the object of interest is offset from the centroid of the view by defining the object of interest within a rectangular bounded area having a horizontal width; and place the object of interest within the rectangular bounded area such that the centroid of the object of interest is horizontally shifted from the boundary of the rectangular bounded area, the more pointed to by the gaze, by a distance corresponding to two-thirds of the horizontal width.

[0086] 17. A non-transitory computer-readable medium storing instructions executable by a processor, the instructions including instructions for performing the following operations: receiving an image data frame from an optical sensor; detecting one or more conference participants within the image data frame; identifying a region of interest for each of the one or more conference participants, wherein identifying the region of interest for each of the one or more conference participants includes: estimating a head pose from a first participant among the one or more conference participants; determining that more of the regions of interest overlap in an overlapping area; detecting an object within the overlapping area; determining that the object within the overlapping area is an object of interest; and presenting a view containing the object of interest in a transmission to a remote endpoint.

[0087] 18. The non-transitory computer-readable medium according to Example 17, wherein the instructions for identifying the region of interest for each of the one or more conference participants further include: instructions for estimating the gaze of a second participant among the one or more conference participants.

[0088] 19. The non-transitory computer-readable medium according to Example 17, wherein the instructions for identifying the region of interest for each of the one or more conference participants further include: instructions for generating a focus map using a neural network.

[0089] 20. The non-transitory computer-readable medium according to Example 17, wherein the instruction for determining that the object within the overlapping region is the object of interest further includes: an instruction for determining that the object corresponds to a person.

[0090] 21. The non-transitory computer-readable medium according to Example 20, wherein the instructions for presenting the view containing the object of interest include instructions for performing the following operations: determining a centroid corresponding to the object of interest; determining the gaze of the object of interest relative to a lens of an optical sensor for capturing the view containing the object of interest, the lens having a central region; determining that the gaze of the object of interest deviates from the normal of the central region by at least fifteen degrees; and positioning the object of interest within the view such that the centroid of the object of interest deviates from the centroid of the view.

[0091] The various examples described above are provided by way of illustration and should not be construed as limiting the scope of this disclosure. Various modifications and changes may be made to the principles and examples described herein without departing from the scope of this disclosure and without departing from the appended claims.

Claims

1. A method for view selection in a teleconference environment, comprising: Receive image data frames from the optical sensor; Detect multiple conference participants within the image data frame; Identifying regions of interest for each of the plurality of meeting participants, wherein identifying the regions of interest for each of the plurality of meeting participants includes: estimating the head pose of a first participant among the plurality of meeting participants; It was determined that most of the regions of interest overlapped in the overlapping region; Detect objects within the overlapping region; Determine that the object within the overlapping region is the object of interest; and Present a view containing the object of interest. Determining that the object within the overlapping area is the object of interest further includes: determining that the object corresponds to a person. Presenting the view containing the object of interest includes: Determine the centroid corresponding to the object of interest; Determine the gaze of the object of interest relative to the lens of the optical sensor, the optical sensor being used to capture the view containing the object of interest, the lens having a central area; Determine that the gaze of the object of interest deviates from the normal of the central region by at least fifteen degrees; and Position the object of interest within the view such that the centroid of the object of interest is offset from the centroid of the view.

2. The method of claim 1, wherein, Identifying the region of interest for each of the plurality of meeting participants also includes estimating the gaze of a second participant among the plurality of meeting participants.

3. The method of claim 2, wherein, The first participant and the second participant are different.

4. The method of claim 1, wherein, Identifying the region of interest for each of the plurality of meeting participants also includes generating a focus map using a neural network.

5. The method of claim 1, wherein, Determining that the object corresponds to a person includes: determining that the object corresponds to a person who has not spoken.

6. The method of claim 1, wherein, Positioning the object of interest within the view such that the centroid of the object of interest is offset from the centroid of the view includes: defining the object of interest within a rectangular bounded area having a horizontal width; and placing the object of interest within the rectangular bounded area such that the centroid of the object of interest is horizontally shifted from the boundary of the rectangular bounded area, which the gaze is pointing towards, by a distance corresponding to two-thirds and one-half of the horizontal width.

7. A teleconference endpoint, comprising: An optical sensor configured to receive image data frames; A processor, which is coupled to the optical sensor, wherein the processor is configured to: Detect multiple conference participants within the image data frame; Regions of interest for each of the plurality of meeting participants are identified by estimating the head pose of the first participant among the plurality of meeting participants; It was determined that most of the regions of interest overlapped in the overlapping region; Detect objects within the overlapping region; Determine that the object within the overlapping region is the object of interest; and Present a view containing the object of interest. The processor is further configured to determine that the object of interest corresponds to a person. The processor is further configured to present the view containing the object of interest in the following manner: Determine the centroid corresponding to the object of interest; Determine the gaze of the object of interest relative to the lens of the optical sensor, the optical sensor being used to capture the view containing the object of interest, the lens having a central area; Determine that the gaze of the object of interest deviates from the normal of the central region by at least fifteen degrees; and Position the object of interest within the view such that the centroid of the object of interest is offset from the centroid of the view.

8. The conference call endpoint of claim 7, wherein, The processor is also configured to: identify the region of interest for each of the plurality of meeting participants, including estimating the gaze of a second participant among the plurality of meeting participants.

9. The conference call endpoint of claim 8, wherein, The first participant and the second participant are different.

10. The conference call endpoint of claim 7, wherein, The processor is also configured to identify the region of interest for each of the plurality of conference participants based on a focus map generated using a neural network.

11. The conference call endpoint of claim 7, wherein, The person in question is not an active speaker.

12. The conference call endpoint of claim 7, wherein, The processor is also configured to: By defining the object of interest within a rectangular bounded area with a horizontal width, the object of interest is positioned within the view such that the centroid of the object of interest is offset from the centroid of the view; and The object of interest is placed within the rectangular bounded area such that the centroid of the object of interest is horizontally shifted from the boundary of the rectangular bounded area, which the gaze is pointing towards, by a distance between three-quarters and one-half of the horizontal width.

13. A non-transitory computer-readable medium storing instructions executable by a processor, the instructions including instructions for performing the following operations: Receive image data frames from the optical sensor; Detect multiple conference participants within the image data frame; Identify regions of interest for each of the plurality of meeting participants, wherein, Identifying the region of interest for each of the plurality of meeting participants includes estimating the head pose of a first participant among the plurality of meeting participants; It was determined that most of the regions of interest overlapped in the overlapping region; Detect objects within the overlapping region; Determine that the object within the overlapping region is the object of interest; and During transmission to the remote endpoint, a view containing the object of interest is presented. The instruction for determining that the object within the overlapping area is the object of interest further includes: an instruction for determining that the object corresponds to a person. The instructions for presenting the view containing the object of interest include instructions for performing the following operations: Determine the centroid corresponding to the object of interest; Determine the gaze of the object of interest relative to the lens of the optical sensor, the optical sensor being used to capture the view containing the object of interest, the lens having a central area; Determine that the gaze of the object of interest deviates from the normal of the central region by at least fifteen degrees; and Position the object of interest within the view such that the centroid of the object of interest is offset from the centroid of the view.

14. The non-transitory computer-readable medium of claim 13, wherein, The instructions for identifying the region of interest for each of the plurality of meeting participants further include instructions for estimating the gaze of a second participant among the plurality of meeting participants.

15. The non-transitory computer-readable medium of claim 13, wherein, The instructions to identify the area of interest for each of the plurality of conference participants further include instructions to generate a heat map using a neural network.