Compositing and scaling angle-separated subscenes

The conference camera system addresses the limitations of existing video conferencing by capturing panoramic scenes and combining sub-scenes to create a seamless single-camera view, enhancing visibility and reducing audio interference for remote participants.

JP7752021B2Active Publication Date: 2025-10-09OWL LABS INC
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
JP2021172415
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2015-04-01
Filing Date
2021-10-21
Publication Date
2025-10-09
Estimated Expiration
2036-04-01

AI Technical Summary

Technical Problem

Existing video conferencing technologies face challenges in capturing and displaying multiple participants within a conference room due to limited camera angles and audio feedback, leading to distorted views and non-verbal cue difficulties for remote parties.

Method used

A conference camera system that captures panoramic video signals with a wide angle and subsamples sub-scene video signals to form a stage scene signal, combining them side-by-side to create a single-camera view, enhancing the visibility of all participants and reducing audio feedback.

Benefits of technology

The system provides a seamless and immersive video experience for remote participants by improving the visibility of all conference attendees and minimizing audio interference, making the in-room experience more natural and comfortable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007752021000001
    Figure 0007752021000001
  • Figure 0007752021000002
    Figure 0007752021000002
  • Figure 0007752021000003
    Figure 0007752021000003
Patent Text Reader

Abstract

An apparatus and method for image capture and enhancement is provided. [Solution] A method by a mobile device (40) for synthesizing angle-separated sub-scenes and / or target sub-scenes within a wide scene collected by the device (100), wherein a densely synthesized single-camera signal is formed from a panoramic video signal captured from a wide camera having an aspect ratio of substantially 2.4:1 or greater. Two or more sub-scene video signals are subsampled in their respective target orientations and synthesized side-by-side to form a stage scene video signal having an aspect ratio of substantially 2:1 or less. At least 80% of the area of ​​the stage scene video signal is subsampled from the panoramic video signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit, pursuant to 35 U.S.C. §119(e), of U.S. Provisional Patent Application Serial No. 62 / 141,822, filed April 1, 2015, the entire disclosure of which is incorporated herein by reference.

[0002] Field Aspects relate to apparatus and methods for image capture and enhancement. [Background technology]

[0003] background Multi-party teleconferencing, video chats, and video conferences often occur with multiple participants in the same conference room connected to at least one remote party.

[0004] In the person-to-person mode of video conferencing software, only one local camera is available, often with a limited horizontal field of view (e.g., 70 degrees). Whether this single camera is positioned in front of one participant or at the head of the table pointed at all participants, it is difficult for the remote party to understand the audio, body language, and non-verbal cues given by those participants in the room who are far from or at an acute angle to the single camera (e.g., seeing a person's outline rather than their face).

[0005] The multi-person mode of video conferencing software adds several different challenges due to the availability of cameras on two or more mobile devices (laptops, tablets, or mobile phones) in the same conference room. The more conference room participants logged into the conference, the greater the audio feedback and crosstalk can be. The camera perspective may be as distant or distorted from the participants as with a single camera. Despite being in the same room, local participants may tend to interact with other participants through their mobile devices (thereby inheriting the same body language and nonverbal cue weaknesses as remote parties).

[0006] There are no known commercial or experimental techniques for compositing, tracking, and / or displaying angle-separated sub-scenes and / or target sub-scenes within a wide scene (e.g., a wide scene of two or more conference participants) in a way that makes setup significantly easier for in-room participants or makes the experience automatically seamless from the perspective of remote participants. Summary of the Invention [Means for solving the problem]

[0007] overview In one aspect of this embodiment, the process for outputting a densely combined single camera signal may record a panoramic video signal having an aspect ratio of substantially 2.4:1 or greater, captured from a wide camera having a horizontal angle of view of substantially 90 degrees or greater. At least two sub-scene video signals may be subsampled from the wide camera at respective target orientations. The two or more sub-scene video signals may be combined side-by-side to form a stage scene video signal having an aspect ratio of substantially 2:1 or less. Optionally, more than 80% of the area of ​​the stage scene video signal may be subsampled from the panoramic video signal. The stage scene video signal may be formatted as a single camera video signal. Optionally, the panoramic video signal has an aspect ratio of substantially 8:1 or greater and is captured from a wide camera having a horizontal field of view of substantially 360 degrees.

[0008] In a related aspect of this embodiment, a conference camera is configured to output a densely combined single-camera signal. An image capture element or wide camera of the conference camera may be configured to capture and / or record a panoramic video signal having an aspect ratio of substantially 2.4:1 or greater, the wide camera having a horizontal angle of view of substantially 90 degrees or greater. A processor operably connected to the image capture element or wide camera may be configured to subsample two or more sub-scene video signals from the wide camera at respective target orientations. The processor may be configured to combine the two or more sub-scene video signals as side-by-side video signals in a memory (e.g., a buffer and / or video memory) to form a stage scene video signal having an aspect ratio of substantially 2:1 or less. The processor may be configured to combine the sub-scene video signals in a memory (e.g., a buffer and / or video memory) such that more than 80% of the area of ​​the stage scene video signal is subsampled from the panoramic video signal. The processor may further be configured to format the stage scene video signal as a single-camera video signal for transport, for example, over USB.

[0009] In any one of the above aspects, the processor may be configured to sub-sample additional sub-scene video signals at respective target orientations from the panoramic video signal, and to combine two or more sub-scene video signals with one or more additional sub-scene video signals to form a stage scene video signal having an aspect ratio of substantially 2:1 or less, including a plurality of side-by-side sub-scene video signals. Optionally, combining the two or more sub-scene video signals with the one or more additional sub-scene video signals to form the stage scene video signal includes transitioning the one or more additional sub-scene video signals into the stage scene video signal by replacing at least one of the two or more sub-scene video signals to form the stage scene video signal having an aspect ratio of substantially 2:1 or less.

[0010] Further optionally, each sub-scene video signal may be assigned a minimum width, and upon completion of its respective transition to the stage scene video signal, each sub-scene video signal may be combined side-by-side at substantially equal to or greater than its minimum width to form the stage scene video signal. Alternatively, or in addition, the combined width of each sub-scene video signal during the transition may increase throughout the transition until the combined width is substantially equal to or greater than its corresponding respective minimum width. Further alternatively, or in addition, the sub-scene video signals may be combined side-by-side at substantially equal to or greater than their minimum width, each combined at a respective width such that the sum of all combined sub-scene video signals is substantially equal to the width of the stage scene video signal.

[0011] In some cases, widths of sub-scene video signals within the stage scene video signal may be combined to vary according to activity criteria detected in one or more target orientations corresponding to the sub-scene video signals, while the width of the stage scene video signal is kept constant. In other cases, combining two or more sub-scene video signals with one or more additional sub-scene video signals to form the stage scene video signal includes transitioning the one or more additional sub-scene video signals into the stage scene video signal by reducing the width of at least one of the two or more sub-scene video signals by an amount corresponding to the width of the one or more additional sub-scene video signals.

[0012] Further optionally, each sub-scene video signal may be assigned a respective minimum width, and each sub-scene video signal may be combined side-by-side at substantially equal to or greater than its corresponding respective minimum width to form a stage sequence. Together with one or more additional sub-scene video signals, at least one of the two or more sub-scene video signals may be transitioned to be removed from the stage scene video signal when the sum of the respective minimum widths of the two or more sub-scene video signals exceeds the width of the stage scene video signal. Optionally, the sub-scene video signal transitioned to be removed from the stage scene video signal corresponds to a respective target orientation for which the activity criterion was most recently met.

[0013] In any of the above aspects, the left-to-right order relative to the wide camera between the respective target orientations of the two or more sub-scene video signals and the one or more additional sub-scene video signals may be preserved when the two or more sub-scene video signals are combined with the one or more additional sub-scene video signals to form the stage scene video signal.

[0014] Further, in any one of the above aspects, each object orientation from the panoramic video signal may be selected depending on a selection criterion detected at each object orientation relative to the wide camera. After the selection criterion is no longer true, the corresponding sub-scene video signal may be transitioned to be removed from the stage scene video signal. Alternatively, or in addition, the selection criterion may include the presence of an activity criterion being met at each object orientation. In this case, the processor may count the time since the activity criterion was met at each object orientation. A predetermined period of time after the activity criterion was met at each object orientation, the corresponding sub-scene signal may be transitioned to be removed from the stage scene video signal.

[0015] In a further variation of the above aspect, the processor may subsample from the panoramic video signal a downscaled panoramic video signal having an aspect ratio of substantially 8:1 or greater, and combine two or more sub-scene video signals with the downscaled panoramic video signal to form a stage scene video signal having an aspect ratio of substantially 2:1 or less, including a plurality of side-by-side sub-scene video signals and a panoramic video signal. Optionally, the two or more sub-scene video signals may be combined with the downscaled panoramic video signal to form a stage scene video signal having an aspect ratio of substantially 2:1 or less, including a plurality of side-by-side sub-scene video signals and a panoramic video signal that is taller than the plurality of side-by-side sub-scene video signals, the panoramic video signal being no more than one-fifth the area of ​​the stage scene video signal and extending substantially across the width of the stage scene video signal.

[0016] In a further variation of the above aspect, the processor or an associated processor may transition the text video signal to the stage scene video signal by subsampling the text video signal from the text document and replacing at least one of the two or more sub-scene video signals with the text video signal.

[0017] Optionally, the processor may set at least one of the two or more sub-scene video signals as a protected sub-scene video signal that is protected from transitioning based on a retention criterion, in which case the processor may transition one or more additional sub-scene video signals to the stage scene video signal by replacing at least one of the two or more sub-scene video signals and / or by transitioning a sub-scene video signal other than the protected sub-scene.

[0018] In some cases, the processor may alternatively or additionally set a sub-scene enhancement operation based on an enhancement criterion, and at least one of the two or more sub-scene video signals is enhanced according to the sub-scene enhancement operation based on the corresponding enhancement criterion. Optionally, the processor may set a sub-scene participant notification operation based on a criterion detected from a sensor, and a local reminder indicator (e.g., a light, blinking, or sound) is activated in response to a corresponding detected criterion. Notification actions are triggered based on the specified criteria.

[0019] In one aspect of this embodiment, a process for tracking a sub-scene at a target orientation within a wide video signal may include monitoring an angular range using an acoustic sensor array and a wide camera observing a field of view of substantially 90 degrees or greater. A first target orientation may be identified along the localization of at least one of the acoustic and visual perceptions detected within the angular range. A first sub-scene video signal may be subsampled from the wide camera along the first target orientation. A width of the first sub-scene video signal may be set according to signal characteristics of at least one of the acoustic and visual perceptions.

[0020] In a related aspect of this embodiment, a conference camera may be configured to output a video signal including a subsampled and scaled subscene from a wide-angle scene and track the subscene and / or target orientation within the wide video signal. The conference camera and / or its processor may be configured to monitor an angular range using an acoustic sensor array and a wide camera observing a field of view of substantially 90 degrees or greater. The processor may be configured to identify a first target orientation along the localization of at least one of acoustic and visual perceptions detected within the angular range. The processor may be further configured to subsample the first subscene video signal from the wide camera along the first target orientation to a memory (buffer or video). The processor may be further configured to set the width of the first subscene video signal according to signal characteristics of at least one of the acoustic and visual perceptions.

[0021] In any of the above aspects, the signal characteristics may represent a confidence level of one or both of the acoustic recognition and the visual recognition. Optionally, the signal characteristics may represent a width of a recognized feature in one or both of the acoustic recognition and the visual recognition. Further optionally, the signal characteristics may correspond to an approximate width of a recognized human face along the first target orientation.

[0022] Alternatively, or in addition, if the width is not set according to the signal characteristics of the visual recognition, the predetermined width may be set along the localization of the detected acoustic recognition within the angular range. Further optionally, the first target orientation may be determined by the visual recognition, and the width of the first sub-scene video signal may then be set according to the signal characteristics of the visual recognition. Further optionally, the first target orientation may be identified as being oriented toward the detected acoustic recognition within the angular range. In this case, the processor may identify a visual recognition proximate to the acoustic recognition, and the width of the first sub-scene video signal may then be set according to the signal characteristics of the visual recognition proximate to the acoustic recognition.

[0023] In another aspect of this embodiment, a processor may be configured to perform a process for tracking sub-scenes at target orientations within a wide video signal, the process including scanning a sub-sampling window through the motion video signal corresponding to a wide camera field of view of substantially 90 degrees or greater. The processor may be configured to identify candidate orientations within the sub-sampling window, each target orientation corresponding to a localization of visual recognition detected within the sub-sampling window. The processor may then record the candidate orientations in a spatial map and may monitor an angular range corresponding to the wide camera field of view using an acoustic sensor array for acoustic recognition.

[0024] Optionally, when an acoustic recognition is detected proximate to one of the candidate orientations recorded in the spatial map, the processor may further snap a first target orientation to substantially correspond to the one of the candidate orientations, and sub-sample a first sub-scene video signal from the wide camera along the first target orientation. Optionally, the processor may be further configured to set a width of the first sub-scene video signal according to a signal characteristic of the acoustic recognition. Further optionally, the signal characteristic may represent a confidence level of the acoustic recognition, or a confidence level of either the acoustic recognition or the visual recognition. The signal characteristics may represent the width of the recognized features within the first target orientation, or both. The signal characteristics may alternatively or additionally correspond to an approximate width of the recognized human face along the first target orientation. Optionally, if the width is not set according to the signal characteristics of the visual recognition, a predetermined width may be set along the localization of the detected acoustic recognition within an angular range.

[0025] In another aspect of this embodiment, a processor may be configured to track a sub-scene in a target orientation, which includes recording a motion video signal corresponding to a wide camera field of view of substantially 90 degrees or greater. The processor may be configured to monitor an angular range corresponding to the wide camera field of view using an acoustic sensor array for acoustic recognition and identify a first target orientation directed toward an acoustic recognition detected within the angular range. A sub-sampling window may be positioned within the motion video signal according to the first target orientation, and a visual recognition may be detected within the sub-sampling window. Optionally, the processor may be configured to sub-sample the first sub-scene video signal captured from the wide camera substantially centered on the visual recognition and set a width of the first sub-scene video signal according to signal characteristics of the visual recognition.

[0026] In a further aspect of this embodiment, a processor may be configured to track a sub-scene at a target orientation within a wide video signal, which includes monitoring an angular range using an acoustic sensor array and a wide camera observing a field of view of substantially 90 degrees or greater. Multiple target orientations may be identified, each oriented toward a localization within the angular range. The processor may be configured to maintain a spatial map of recorded characteristics corresponding to the target orientations and to subsample the sub-scene video signal from the wide camera substantially along one or more target orientations. A width of the sub-scene video signal may be set according to the recorded characteristics corresponding to at least one target orientation.

[0027] In a further aspect of this embodiment, the processor may be configured to execute a process for tracking sub-scenes at target orientations within a wide video signal, the process including monitoring an angular range using an acoustic sensor array and a wide camera observing a field of view of substantially 90 degrees or greater, and identifying multiple target orientations, each oriented toward a localization within the angular range. A sub-scene video signal may be sampled from the wide camera substantially along at least one target orientation, and the width of the sub-scene video signal may be set by expanding the sub-scene video signal until a threshold based on at least one recognition criterion is met. Optionally, a change vector for each target orientation may be predicted based on a change in one of the speed and direction of the recorded characteristic corresponding to the localization, and the position of the target orientation may be updated based on the prediction. Optionally, a search region for the localization may be predicted based on the most recent position of the recorded characteristic corresponding to the localization, and the position of the localization may be updated based on the prediction. [Brief explanation of the drawings]

[0028] [Figure 1A] 1 is a schematic block diagram of an embodiment of a device suitable for synthesizing, tracking, and / or displaying angle-separated sub-scenes and / or target sub-scenes within a wide scene collected by device 100. FIG. [Figure 1B] 1 is a schematic block diagram of an embodiment of a device suitable for synthesizing, tracking, and / or displaying angle-separated sub-scenes and / or target sub-scenes within a wide scene collected by device 100. FIG. [Figure 2A] 1A and 1B is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing wide and / or panoramic scenes for the device 100 of FIGS. 1A and 1B. [Figure 2B] 1A and 1B is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing wide and / or panoramic scenes for the device 100 of FIGS. 1A and 1B. [Figure 2C]1A and 1B is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing wide and / or panoramic scenes for the device 100 of FIGS. 1A and 1B. [Figure 2D] 1A and 1B is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing wide and / or panoramic scenes for the device 100 of FIGS. 1A and 1B. [Figure 2E] 1A and 1B is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing wide and / or panoramic scenes for the device 100 of FIGS. 1A and 1B. [Figure 2F] 1A and 1B is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing wide and / or panoramic scenes for the device 100 of FIGS. 1A and 1B. [Figure 2G] 1A and 1B is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing wide and / or panoramic scenes for the device 100 of FIGS. 1A and 1B. [Figure 2H] 1A and 1B is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing wide and / or panoramic scenes for the device 100 of FIGS. 1A and 1B. [Figure 2J] 1A and 1B is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing wide and / or panoramic scenes for the device 100 of FIGS. 1A and 1B. [Figure 2K] 1A and 1B is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing wide and / or panoramic scenes for the device 100 of FIGS. 1A and 1B. [Figure 2L] 1A and 1B is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing wide and / or panoramic scenes for the device 100 of FIGS. 1A and 1B. [Figure 3A]A top-down view of a conference camera use case showing three participants. [Figure 3B] FIG. 1 is a top-down view of a conference camera panoramic image signal showing three participants. [Figure 4A] A top-down view of a conference camera use case showing a conference table with three participants and including depictions of face width settings or sub-scene identification. [Figure 4B] FIG. 10 is a top-down view of a conference camera panoramic image signal showing three participants and including depictions of face width settings or sub-scene identification. [Figure 5A] A top-down view of a conference camera use case showing a conference table with three participants, including depictions of shoulder spacing or sub-scene identification. [Figure 5B] FIG. 10 is a top-down view of a conference camera panoramic image signal showing three participants and including depictions of shoulder spacing or sub-scene identification. [Figure 6A] A top-down view of a conference camera use case showing a conference table with three participants and a whiteboard, including a depiction of a wider sub-scene identification. [Figure 6B] FIG. 1 is a top-down view of a conference camera panoramic image signal showing three participants and a whiteboard, including depictions of identification of wider sub-scenes. [Figure 7A] FIG. 1 is a top-down view of a conference camera use case showing a 10-seat conference table showing five participants and including depictions of visual and acoustic minimum width and orientation identification. [Figure 7B] 1 is a top-down view of a conference camera panoramic image signal showing five participants and including depictions of visual and acoustic minimum width and orientation identification. [Figure 8A] 1 is a schematic diagram of a conference camera video signal, a minimum width, and extraction of sub-scene and panoramic video signals to be composited into a stage scene video signal; [Figure 8B]1 is a schematic diagram of a sub-scene video signal and a panoramic video signal to be composited into a stage scene video signal. [Figure 8C] FIG. 1 illustrates a possible composite output or stage scene video signal. [Figure 8D] FIG. 1 illustrates a possible composite output or stage scene video signal. [Figure 8E] FIG. 1 illustrates a possible composite output or stage scene video signal. [Figure 9A] 1 is a schematic diagram of a conference camera video signal, a minimum width, and extraction of alternative sub-scene and alternative panoramic video signals to be composited into a stage scene video signal; [Figure 9B] 1 is a schematic diagram of an alternative sub-scene video signal and an alternative panoramic video signal to be composited into a stage scene video signal; [Figure 9C] FIG. 10 illustrates a possible alternative composite output or stage scene video signal. [Figure 9D] FIG. 10 illustrates a possible alternative composite output or stage scene video signal. [Figure 9E] FIG. 10 illustrates a possible alternative composite output or stage scene video signal. [Figure 9F] 1 is a schematic diagram of a panoramic video signal adjusted so that the conference table images are arranged for a more natural and comfortable view. [Figure 10A] FIG. 1 is a schematic diagram of a possible composite output or stage scene video signal. [Figure 10B] FIG. 1 is a schematic diagram of a possible composite output or stage scene video signal. [Figure 11A] 10 is a schematic diagram of an alternative way in which the videoconferencing software may display a composite output or stage scene video signal. [Figure 11B] 10 is a schematic diagram of an alternative way in which the videoconferencing software may display a composite output or stage scene video signal. [Figure 12] FIG. 1 is a flowchart including steps for compositing a stage scene (video signal) video signal. [Figure 13] FIG. 10 is a detailed flowchart including steps for synthesizing and generating a sub-scene (sub-scene video signal) based on target orientation. [Figure 14] FIG. 10 is a detailed flowchart including steps for compositing a sub-scene into a stage scene video signal. [Figure 15] FIG. 10 is a detailed flowchart including steps for outputting a composited stage scene video signal as a single camera signal. [Figure 16] FIG. 10 is a detailed flowchart illustrating a first mode of performing steps for setting the localization and / or target orientation and / or sub-scene width. [Figure 17] FIG. 10 is a detailed flowchart including a second mode of executing steps for setting the localization and / or target orientation and / or sub-scene width. [Figure 18] FIG. 10 is a detailed flowchart illustrating a third mode of performing steps for setting the localization and / or target orientation and / or sub-scene width. [Figure 19] 3A-5B , which illustrate the operation of an embodiment including a conference camera attached to a local PC with a videoconferencing client receiving a single camera signal, the PC then being connected to the Internet, and two remote PCs, etc. also receiving the single camera signal within the videoconferencing display. [Figure 20] 3A-5B , which illustrate the operation of an embodiment including a conference camera attached to a local PC with a videoconferencing client receiving a single camera signal, the PC then being connected to the Internet, and two remote PCs, etc. also receiving the single camera signal within the videoconferencing display. [Figure 21]3A-5B , which illustrate the operation of an embodiment including a conference camera attached to a local PC with a videoconferencing client receiving a single camera signal, the PC then being connected to the Internet, and two remote PCs, etc. also receiving the single camera signal within the videoconferencing display. [Figure 22] FIG. 22 illustrates a variation of the system of FIGS. 19-21 in which the videoconferencing clients use overlapping video views instead of separate adjacent views. [Figure 23] 22A-22C show a variation of the system of FIGS. 19-21, substantially corresponding to FIGS. 6A-6B, including a high resolution camera view for the whiteboard. [Figure 24] 22A-22D show variations of the systems of FIGS. 19-21 that include high-resolution text document views (eg, text editors, word processing, presentations, or spreadsheets). [Figure 25] FIG. 1C is a schematic diagram of an arrangement in which a videoconferencing client is instantiated for each subscene, using an arrangement similar to that of FIG. 1B. [Figure 26] FIG. 27 is a schematic diagram of some exemplary icons and symbols used throughout FIGS. 1-26. DETAILED DESCRIPTION OF THE INVENTION

[0029] Detailed Description Conference Camera 1A and 1B are schematic block diagrams of embodiments of a device suitable for synthesizing, tracking, and / or displaying angle-separated sub-scenes and / or target sub-scenes within a wide scene collected by a device that is a conference camera 100.

[0030] FIG. 1A illustrates a device configured as a conference camera 100 or conference “webcam,” e.g., to communicate as a USB peripheral connected to a USB host or hub of a connected laptop, tablet, or mobile device 40, and to provide a single video image with an aspect ratio, pixel count, and proportions commonly used by off-the-shelf video chat or video conferencing software such as Google Hangouts, Skype, or Facetime. Device 100 includes a “wide camera” 2, 3, or 5, e.g., a camera oriented to overlook a conference of attendees or participants M1, M2, ... Mn, capable of capturing two or more attendees. Camera 2, 3, or 5 may include one digital imaging device or lens, or two or more digital imaging devices or lenses (e.g., stitched in software or otherwise). It should be noted that, depending on the location of device 100 within the conference, the field of view of wide camera 2, 3, or 5 may be 70 degrees or less. However, in one or more embodiments, wide cameras 2, 3, 5 are useful in the center of a meeting, where the wide cameras may have a horizontal field of view that is substantially 90 degrees, or greater than 140 degrees (not necessarily continuous), or up to 360 degrees.

[0031] In large conference rooms (e.g., those designed to accommodate eight or more people), it may be useful to have multiple wide-angle camera devices that record a wide field of view (e.g., substantially 90 degrees or more) and collectively stitch together the very wide scene to capture the most pleasing angle. For example, a wide-angle camera at the far end of a long (10'-20') table may provide an unsatisfactory distant view of speaker SPKR, while having multiple cameras distributed across the table (e.g., one for every five seats) may provide at least one satisfactory or pleasing view. Cameras 2, 3, and 5 may capture or record a panoramic scene (e.g., in an H:V horizontal-vertical ratio, e.g., an aspect ratio of 2.4:1 to 10:1) and / or make this signal available via a USB connection.

[0032] As described with respect to FIGS. 2A-2L, the height of the wide cameras 2, 3, and 5 from the base of the conference camera 100 is preferably 8 inches. (20.32 cm) Because the cameras 2, 3, and 5 are larger than typical laptop screens in a conference, they may have unobstructed and / or near-eye-level views of conference attendees M1, M2, ... Mn. Microphone array 4 includes at least two microphones and can obtain target orientation to nearby sounds or speech by beamforming, relative time-of-flight, localization, or received signal strength differences, as known in the art. Microphone array 4 may include multiple microphone pairs oriented to cover at least substantially the same angular range as the field of view of wide camera 2.

[0033] Microphone Array 4 is an 8-inch (20.32 cm)Optionally, the microphone array 4 is arranged with wide cameras 2, 3, 5 at a height higher than the microphone array 4, so that there is still a direct "line of sight" between the array 4 and attendees M1, M2, ... Mn when they are speaking, unobstructed by a typical laptop screen. A CPU and / or GPU (and associated circuitry, such as camera circuitry) 6 for processing computations and graphical events is connected to each of the wide cameras 2, 3, 5 and the microphone array 4. ROM and RAM 8 are connected to the CPU and GPU 6 for holding and receiving executable code. A network interface and stack 10 is provided for USB, Ethernet, and / or WiFi connected to the CPU 6. One or more serial buses interconnect these electronic components, which may be powered by DC, AC, or battery power.

[0034] The camera circuitry of cameras 2, 3, and 5 may output processed or rendered images or video streams as single-camera image signals, video signals, or streams in a landscape orientation with an "H:V" horizontal-vertical ratio or aspect ratio of 1.25:1 to 2.4:1 or 2.5:1 (including ratios of 4:3, 16:10, and 16:9, for example), and / or, using suitable lenses and / or stitching circuitry, as described above, as a panoramic image or video stream as a single-camera image signal of substantially 2.4:1 or greater. The conference camera 100 of FIG. 1A may be connected, typically as a USB peripheral, to a laptop, tablet, or mobile device 40 (having a display, network interface, computing processor, memory, camera, and microphone components interconnected by at least one bus) hosting multi-party teleconferencing, videoconferencing, or videochat software and connectable for teleconferencing with remote clients 50 over the Internet 60.

[0035] FIG. 1B is a variation of FIG. 1A in which both the device 100 and the videoconferencing device 40 of FIG. 1A are integrated. The camera circuitry output as a single camera image signal, video signal, or video stream is directly available to the CPU, GPU, associated circuitry, and memory 5, 6, and the videoconferencing software is instead hosted by the CPU, GPU, and associated circuitry and memory 5, 6. The device 100 can be directly connected to the Internet 60 or INET (e.g., via WiFi or Ethernet) for videoconferencing with a remote client 50. The display 12 provides a user interface for operating the videoconferencing software and for presenting the videoconferencing views and graphics described herein to conferencing participants M1, M2, ... M3. The device or conference camera 100 of FIG. 1A may instead be directly connected to the Internet 60, thereby enabling the remote client 50 to record video directly to a remote server or to access video live from such a server.

[0036] 2A-2L are schematic diagrams of embodiments of conference camera 14 or camera tower 14 arrangements suitable for capturing wide and / or panoramic scenes for the device or conference camera 100 of FIGS. 1A and 1B. While "camera tower" 14 and "conference camera" 14 may be used substantially interchangeably herein, a conference camera need not be a camera tower. The height of the wide cameras 2, 3, 5 from the base of the device 100 in FIGS. 2A-2L is preferably 8 inches. (20.32 cm) Larger than 15 inches (38.1 cm) is smaller than.

[0037] In the camera tower 14 arrangement of FIG. 2A, multiple cameras are mounted at camera level (8 to 15 inches) on the camera tower 14. (20.32 to 38.1 cm)) are arranged around the periphery and equiangularly spaced. The number of cameras is determined by the camera's field of view and the angle to be spanned, and the cumulative angle spanned should have overlap between the individual cameras when forming a panoramic stitched view. For example, in FIG. 2A, four cameras 2a, 2b, 2c, and 2d (labeled 2a-2d), each with a 100-110 degree field of view (shown in dashed lines), are arranged at 90 degrees from each other to provide a 360-degree cumulative or stitchable or stitched view around the camera tower 14.

[0038] For example, in Figure 2B, three cameras 2a, 2b, and 2c (labeled 2a-2c), each with a field of view of 130 degrees or more (shown in dashed lines), are arranged at 120 degrees from one another to again provide a cumulative or stitchable 360-degree view around tower 14. The vertical fields of view of cameras 2a-2d are smaller than their horizontal fields of view, e.g., less than 80 degrees. The images, videos, or sub-scenes from each camera 2a-2d may be processed before or after known optical corrections, such as stitching, dewarping, or distortion compensation, to identify target orientations or sub-scenes, but will typically be so corrected before output.

[0039] In the camera tower 14 arrangement of FIG. 2C, a single fisheye or near-fisheye camera 3a oriented upward is positioned at camera level (8 to 15 inches) of the camera tower 14. (20.32 to 38.1 cm) ) on top of the conference table. In this case, the fisheye camera lens is arranged with a 360-degree continuous horizontal view and a vertical field of view of approximately 215 (e.g., 190-230) degrees (shown in dashed lines). Alternatively, a single catadioptric "cylindrical imaging" camera or lens 3b, e.g., having a cylindrical transmissive shell, upper parabolic mirror, black central post, and telecentric lens configuration as shown in FIG. 2D, is arranged with a 360-degree continuous horizontal view and a vertical field of view of approximately 40-80 degrees, and is centered approximately on the horizon. For each of the fisheye and cylindrical imaging cameras, the camera is positioned 8-15 inches from the conference table. (20.32~38.1cm)The upwardly positioned vertical field of view extends below the horizon, allowing attendees M1, M2...Mn around the conference table to be imaged down to waist height or below. The images, videos or sub-scenes from each camera 3a or 3b may be processed before or after known optical corrections for fisheye or catadioptric lenses, such as dewarping or distortion compensation, to identify target orientations or sub-scenes, but will typically be so corrected before output.

[0040] In the camera tower 14 arrangement of Figure 2L, multiple cameras are mounted at camera level (8 to 15 inches) on the camera tower 14. (20.32 to 38.1 cm) ) and are equiangularly spaced. The number of cameras in this case is not intended to form a completely continuous, stitched panoramic view, and the cumulative angles spanned do not have overlap between the individual cameras. For example, in FIG. 2L, two cameras 2a and 2b, each with a field of view of 130 degrees or more (shown by dashed lines), are arranged at 90 degrees from each other to provide separate views encompassing approximately 260 degrees or more on either side of the camera tower 14. This arrangement may be useful for a long conference table CT. For example, in FIG. 2E, two cameras 2a-2b are panning and / or rotatable about a vertical axis to cover the target orientations B1, B2...Bn described herein. Images, videos, or sub-scenes from each camera 2a-2b may be scanned or analyzed as described herein before or after optical correction.

[0041] 2F and 2G show head and end table arrangements, i.e., each of the camera towers 14 shown in FIGS. 2F and 2G is intended to be advantageously placed at the head of a conference table CT. As shown in FIGS. 3A-6A, a large flat panel display FP for presentations and video conferencing is often placed at the head or end of a conference table CT, and the arrangements of FIGS. 2F and 2G are instead intended to be directly in front of the flat panel FP. In the camera tower 14 arrangement of FIG. 2F, two cameras with approximately 130-degree fields of view are positioned at 120 degrees from each other, covering two sides of a long conference table CT. A display and touch interface 12 is oriented above the table (particularly useful when there are no flat panels FP on the walls) to display the client for the videoconferencing software. This display 12 can be a connected, connectable, or detachable tablet or mobile device. In the camera tower arrangement of FIG. 2G, one high-resolution, arbitrarily tilting camera 7 (optionally connected to its own independent videoconferencing client software or instance) can be directed toward the target object (such as a whiteboard WB or a page or paper on the table CT surface), and two independently panning and / or tilting cameras 5a, 5b with, for example, 100-110-degree fields of view are oriented or oriented to cover the target orientation.

[0042] Images, videos, or sub-scenes from each camera 2a, 2b, 5a, 5b, 7 can be scanned or analyzed as described herein before or after optical correction. Figure 2H shows a variation in which two identical units, each with two 100-130-degree cameras 2a-2b or 2c-2d arranged 90 degrees apart, can be used independently as a >180-degree view unit at the head or end of a table CT, but can also optionally be combined back-to-back to form a unit substantially identical to the unit of Figure 2A, with four cameras 2a-2d appropriately placed in the center of a conference table CT spanning the entire room. Each of the tower units 14, 14 in Figure 2H would be provided with network and / or physical interfaces to form the combined unit. The two units may alternatively or additionally be arranged freely or coordinately as described below with respect to Figures 2K, 6A, 6B, and 14.

[0043] In FIG. 2J, a fisheye camera or lens 3a similar to the camera in FIG. 2c (physically and / or conceptually interchangeable with catadioptric lens 3b) is mounted at camera level (8 to 15 inches) on camera tower 14. (20.32 to 38.1 cm) ) on top of a videoconferencing system. One rotatable, high-resolution, arbitrarily tilted camera 7 (optionally connected to its own independent videoconferencing client software or instance) can be directed toward an object of interest (such as a page or paper on a whiteboard WB or table CT surface). As shown in FIGS. 6A, 6B, and 14, this arrangement works advantageously when a first videoconferencing client (on or connected to a “conference room (local) display” in FIG. 14) receives the composited sub-scene from scene SC cameras 3a, 3b as a single camera image or composite output CO, e.g., via a first physical or virtual network interface or channel 10a, and a second videoconferencing client (residing in device 100 in FIG. 14 and connected to the Internet via a second physical or virtual network interface or channel 10b) receives the independent high-resolution image from camera 7.

[0044] FIG. 2K shows a similar arrangement in which separate video conferencing channels for images from cameras 3a, 3b, and 7 may also be advantageous, but in the arrangement of FIG. 2K, each camera 3a, 3b, and 7 has its own tower 14, optionally connected to the rest of the towers 14 via an interface 15 (which may be wired or wireless). In the arrangement of FIG. 2K, a panoramic tower 14 with scene SC cameras 3a, 3b may be placed in the center of the conference table CT, and a directed high-resolution tower 14 may be placed at the head of the table CT, or anywhere a directed, high-resolution, separate client image or video stream is of interest. Images, videos, or sub-scenes from each camera 3a, 7 may be scanned or analyzed as described herein before or after optical correction.

[0045] Using the conference camera 3A, 3B and 12, according to an embodiment of the method for synthesizing and outputting a photographed scene, a device or conference camera 100 (or 200) is placed on, for example, a circular or rectangular conference table CT. The device 100 may be positioned according to the convenience or intention of the conference participants M1, M2, M3...Mn.

[0046] In any typical conference, participants M1, M2... Mn will be angularly dispersed relative to device 100. If device 100 is placed in the middle of participants M1, M2... Mn, the participants may be captured with a panoramic camera, as described herein. Conversely, if device 100 is placed to one side of the participants (e.g., at one end of a table or mounted on a flat panel FP), a wide camera (e.g., 90 degrees or more) may be sufficient to span participants M1, M2... Mn.

[0047] As shown in FIG. 3A, each of participants M1, M2... Mn will have a respective orientation B1, B2... Bn from device 100, measured, for example, for illustrative purposes, from origin OR. Each orientation B1, B2... Bn can be a range of angles or a nominal angle. As shown in FIG. 3B, an "unrolled," projected, or dewarped fisheye, panoramic, or wide scene SC includes an image of each participant M1, M2... Mn aligned at their expected respective orientation B1, B2... Bn. In particular, in the case of a rectangular table CT and / or alignment of device 100 on one side of the table CT, the image of each participant M1, M2... Mn may be foreshortened or include perspective distortion according to the participant's facing angle (schematically depicted with an expected foreshortening direction in FIG. 3B and throughout the figures). Perspective and / or visual geometric corrections, as known to those skilled in the art, may be applied to images, sub-scenes, or scenes SC that have foreshortening or perspective distortion, but may not be necessary.

[0048] Face detection and widening As an example, modern face detection libraries and APIs (out of over 50 available APIs and SDKs, e.g., Android's FaceDetector) use common algorithms. Face class, Objective C CIDetector class and CIFaceFeature object, OpenCV CascadeClassifier class using Haar cascade) usually use interpupillary distance, etc. and the spatial location of facial features and facial pose. If participant Mn's ears are to be included in the range, a rough lower bound for face width estimation may be about 2 times the interpupillary distance / angle, and a rough upper bound may be 3 times the interpupillary distance / angle. A rough lower bound for portrait width estimation (i.e., head plus some shoulder width) may be 2 times the face width / angle, and a rough upper bound may be 4 times the face width / angle. Alternatively, a fixed angle or other more direct setting of the sub-scene width may be used.

[0049] 4A-4B and 5A-5B illustrate one exemplary two-stage and / or separate identification of both face width and shoulder width (either one of which may be a minimum width as described herein for setting the initial subscene width). As shown in FIGS. 4A and 4B, face widths FW1, FW2...FWn are obtained from a panoramic scene SC, set according to interpupillary distance or other dimensional analysis of facial features (features, classes, colors, segments, patches, textures, trained classifiers, or other characteristics). In contrast, in FIGS. 5A, 5B, 6A, and 6B, shoulder widths SW1, SW2...SWn are set according to the same analysis and scaled by approximately three or four times or according to a default audio decomposition or width.

[0050] Compositing angle-separated subscenes 7A and 7B show a use case of the conference camera 100, showing five participants M1, M2, M3, M4, and M5, a conference table CT seating approximately 10 people, and a conference camera panoramic image signal SC, respectively, including depictions of the identification of a visual minimum width Min. 2 and a corresponding angular range target orientation B5 and an acoustic minimum width Min. 5 and a corresponding vector target orientation B2. This is a view looking down from above.

[0051] 7A, the conference camera 100 is positioned in the center of a long conference table CT that seats 10. Thus, participants M1, M2, and M3 at the center of the table CT are least foreshortened and occupy the largest image area and angular view of the camera 100, while participants M5 and M4 at the ends of the table CT are most foreshortened and occupy the smallest image area.

[0052] In Figure 7B, the entire scene video signal SC is, for example, a 360-degree video signal, and includes all of the participants M1...M5. The conference table CT appears in the scene SC with a distorted "W" shape characteristic of a panoramic view, while the participants M1...M5 appear at different sizes and with different foreshortened aspects depending on their position and distance from the conference camera 100 (simply represented schematically with a rectangular body and an oval head). As shown in Figures 7A and 7B, each participant M1...M5 is represented in memory 8 by their respective orientation B1...B5, which may be determined by acoustic or visual or sensor localization of sounds, actions, or features. As depicted in Figures 7A and 7B, participant M2 may be localized by face detection (and a corresponding vector-like orientation B2 and minimum width Min.2 stored in memory and determined proportionally to the face width obtained from the face detection heuristic), and participant M5 may be localized by beamforming, relative signal strength, and / or time-of-flight of speech-like audio signals (and a corresponding sector-like orientation B5 and minimum width Min.5 stored in memory and determined proportionally to the estimated resolution of acoustic array 4).

[0053] FIG. 8A shows a schematic diagram of the conference camera 100 video signal, the minimum width Min. n, and the extraction of sub-scene video signals SS2, SS5, and panoramic video signal SC.R to be composited into the stage scene video signal STG, CO. The upper part of FIG. 8A essentially reproduces FIG. 7B. As shown in FIG. 8A, the entire scene video signal SC from FIG. 7B can be subsampled according to target orientation (limited to orientations B2 and B5 in this example) and width (limited to widths Min. 2 and Min. 5 in this example). The sub-scene video signal SS2 is at least as wide as the (visually determined) face width limit Min. 2, but may be wider or scaled wider relative to the width, height, and / or available area of ​​the stage STG, or the aspect ratio and available area of ​​the composite output CO. The sub-scene video signal SS5 is at least as wide as the (acoustically determined) acoustic estimate Min. 5, but may be wider or scaled wider and limited as well. The scaled-down panoramic scene SC.R in this capture is a top-and-bottom cropped version of the entire scene SC, in this case cropped to an aspect ratio of 10:1. Alternatively, the scaled-down panoramic scene SC.R may be derived from the entire panoramic scene video signal SC by proportional or anamorphic scaling (e.g., the top and bottom remain but are more compressed than the center). In either case, in the example of Figures 8A and 8B, three different video signal sources SS2, SS5, and SC.R are available to be composited onto stage STG or composite output CO.

[0054] Figure 8B essentially reproduces the lower part of Figure 8A and shows a schematic diagram of the sub-scene video signals SS2, SS5 and panoramic video signal SC.R to be composited into the stage scene video signal STG or CO. Figures 8C to 8E show three possible composite outputs or stage scene video signals STG or CO.

[0055] In the composite output CO or stage scene video signal STG shown in Figure 8C, the scaled down panoramic video signal SC.R is composited across the entire top of the stage STG, in this case occupying less than 1 / 5 or 20% of the stage area. Subscene SS1 has also been composited to occupy its minimum area, not scaled overall, but stretched to fill approximately half the stage width. Subscene SS2 has also been composited to occupy at least its (much smaller) minimum area, not scaled overall, but stretched to fill approximately half the stage width. In this composite output CO, the two subscenes are given approximately the same area, but the participants are of different apparent sizes corresponding to their distance from camera 100. It should also be noted that the left-right or clockwise order of the two composited subscenes is the same as the order of the participants in the room or target orientations from camera 100 (and as they appear in the scaled-down panoramic view SC.R). Furthermore, any of the transitions described herein may be used in compositing subscene video signals SS2 and SS5 into the stage video signal STG. For example, both sub-scenes may simply fill the stage STG instantly, or one may slide in from its corresponding left or right stage direction to fill the entire stage, and then the other may slide in from its corresponding left or right stage direction to gradually narrow, etc., in either case with the sub-scene window, frame, outline, etc. displaying the video stream throughout the transition.

[0056] In the composite output CO or stage scene video signal STG shown in FIG. 8D, the scaled-down panoramic video signal SC.R is similarly composited into the scene STG, but each of signals SS5 and SS2 has been proportionally scaled or zoomed so that participants M5 and M2 occupy a larger area of ​​the stage STG. The minimum width of each signal SS5 and SS2 is also depicted zoomed in, so that signals SS5 and SS2 still occupy more than their respective minimum widths, but each has been widened to fill approximately half of the stage (in the case of SS5, the minimum width is half the stage). Participants M5 and M3 are substantially equal in size on the stage STG, or in the composite output signal CO.

[0057] In the composite output CO or stage scene video signal STG shown in FIG. 8E, the scaled-down panoramic video signal SC.R is similarly composited into the scene STG, but signals SS5 and SS2 are scaled or zoomed accordingly. While sub-scene signals SS5 and SS2 still occupy more than their respective minimum widths, they are each widened to fill different portions of the stage. In this case, sub-scene signal SS5 is not scaled up or zoomed, but has a wider minimum width, occupying more than two-thirds of the area of ​​the stage SG. Meanwhile, signal SS2's minimum width is depicted zoomed, occupying approximately three times its minimum width. One situation in which the relative proportions and conditions of FIG. 8E may arise is when participant M5 is not visually localized, given a widely uncertain (low confidence level) target orientation and a wide minimum width, and further, when participant M5 continues speaking for a long period of time, arbitrarily increasing the occupancy of sub-scene SS5 on the stage STG. At the same time, participant M2 may have reliable face width detection, allowing sub-scene SS2 to be scaled and / or widened to consume an area larger than its minimum width.

[0058] FIG. 9A also shows a schematic diagram of the conference camera 100 video signal, a minimum width Min.n, and the extraction of alternative sub-scene video signals SSn and alternative panoramic video signals SC.R to be composited into the stage scene video signal. The upper part of FIG. 9A essentially replicates FIG. 7B, except that participant M1 is the current speaker and the corresponding sub-scene SS1 has a corresponding minimum width Min.1. As shown in FIG. 9A, the entire scene video signal SC from FIG. 7B can be subsampled according to target orientations (here orientations B1, B2, and B5) and widths (here widths Min.1, Min.2, and Min.5). Each of the sub-scene video signals SS1, SS2, and SS5 is at least as wide as its respective minimum width Min.1, Min.2, and Min.5 (visually, acoustically, or sensor-determined), but may be wider relative to the width, height, and / or available area of ​​the stage STG or the aspect ratio and available area of ​​the composite output CO. It may be scaled wider. The scaled-down panoramic scene SC.R in this capture is a top / bottom and side cropped version of the entire scene SC, in this case with an aspect ratio of approximately 7.5:1, cropped to span only the most relevant / proximate speakers M1, M2, and M5. In the example of Figures 9A and 9B, four different video signal sources SS1, SS2, SS5, and SC.R are available to be composited onto the stage STG or composite output CO.

[0059] Figure 9B essentially reproduces the bottom portion of Figure 9A and shows a schematic diagram of the sub-scene and panoramic video signals to be composited into the stage scene video signal. Figures 9C to 9E show three possible composite outputs or stage scene video signals.

[0060] In the composite output CO or stage scene video signal STG shown in FIG. 9C, the scaled-down panoramic video signal SC.R has been composited almost completely across the top of the stage STG, in this case occupying less than one-quarter of the stage area. Subscene SS5 has again been composited to occupy at least its minimum area, not scaled overall, but stretched to fill approximately one-third of the stage width. Subscenes SS2 and SS1 have also been composited to occupy at least their smaller minimum areas, not scaled overall, and each stretched to fill approximately one-third of the stage width. In this composite output CO, the three subscenes are given approximately the same area, but the participants are of different apparent sizes corresponding to their distance from camera 100. The left-right or clockwise order of the two composited or shifted subscenes remains the same as the order of the participants in the room or target orientations from camera 100 (and as they appear in the scaled-down panoramic view SC.R). Additionally, any of the transitions described herein may be used in compositing the sub-scene video signals SS1, SS2, SS5 into the stage video signal STG. In particular, the transitions may be more comfortable as sliding transitions that approach in or from the same left-right order as the scaled-down panoramic view SC.R (e.g., if M1 and M2 are already on stage, M5 slides in from stage right; if M1 and M5 are already on stage, M2 slides in between them from above or below; if M2 and M5 are already on stage, M1 slides in from stage left, preserving the order of M1, M2, M5 in the panoramic view SC.R).

[0061] In the composite output CO or stage scene video signal STG shown in FIG. 9D, the scaled-down panoramic video signal SC.R is similarly composited into the scene STG, but each of the signals SS1, SS2, and SS5 has been proportionally scaled or zoomed so that participants M1, M2, and M5 occupy a larger area of ​​the stage STG. The minimum width of each signal SS1, SS2, and SS5 is also depicted zoomed, so that signals SS1, SS2, and SS5 still occupy at least their respective zoomed minimum widths, but sub-scene SS5 has been widened to fill an area on the stage slightly larger than its zoomed minimum width, with SS5 occupying 60 percent of the stage width, SS2 occupying only 15 percent, and SS3 occupying the remaining 25 percent. Participants M1, M2, and M5 are of substantially equal height or face size on the stage STG or in the composite output signal CO, but participant M2 and sub-scene SS2 may be substantially cropped to show only a head and / or slightly larger than body width.

[0062] In the composite output CO or stage scene video signal STG shown in FIG. 9E, the scaled down panoramic video signal SC.R is similarly composited into the scene STG, but each of the signals SS1, SS2, and SS5 has been scaled or zoomed accordingly. While the sub-scene signals SS1, SS2, and SS5 still occupy more than their respective minimum widths, each has been widened to fill a different amount of the stage. In this case, none of the sub-scene signals SS1, SS2, and SS5 has been scaled up or zoomed, but rather the nearest or most relevant sub-scene signals SS1, SS2, and SS5 have been scaled up or zoomed. Subscene SS1, with speaker M1 present, occupies more than half of stage SG. Meanwhile, subscenes SS2 and SS5 each occupy a smaller or reduced portion of stage STG, but because subscene SS5 has the smallest width, the further reduction in occupancy of stage STG is taken from subscene SS2 or SS1. One situation in which the relative proportions and conditions of Figure 9E arise may be when visual localization can be achieved for participant M1, but participant M1 continues to speak for an extended period of time, arbitrarily increasing the occupancy of subscene SS1 of stage STG relative to the other two subscenes.

[0063] In the panoramic scene SC or scaled-down panoramic scene SC.R depicted in FIG. 9F , the conference camera 1000 is not centered on the table CT but is instead placed at one end of the table CT (e.g., as shown by the dashed position on the right side of FIG. 7A ), and the flat panel FP represents the remote conference participants. In this case, the conference table CT still appears as a highly distorted “W” shape. As shown in the top of FIG. 9F , if the index direction or origin OR of the conference camera 100 or the panoramic scene SC is oriented so that the edge of the high-aspect ratio panoramic scene SC “cuts” the conference table CT, it is very difficult to see the positions of people around the table CT. However, if the index direction or origin OR of the conference camera 100 or the panoramic scene is arranged so that the table CT is continuous and / or everyone is positioned facing one side, the scene becomes more natural. According to this embodiment, the processor 6 may perform image analysis to change the index position or origin position of the panoramic image. In one example, the index or origin position of the panoramic image may be "rotated" so that the area of ​​a single contiguous segmentation of the image patch corresponding to the table region is maximized (e.g., the table does not crack). In another example, the index or origin position of the panoramic image may be "rotated" so that the two closest or largest facial recognitions are furthest from each other (e.g., the table does not crack). In a third example, the index or origin position of the panoramic image may be "rotated" so that the lowest height segmentation of the image patch corresponding to the table region is located at the panorama edge (e.g., the "W" shape is rotated to place the table edge closest to the conference camera 100 at the panorama edge).

[0064] Figure 10A shows a schematic diagram of a possible composite output CO or stage scene video signal STG, substantially replicating the composite output signal CO or stage video signal STG of Figure 9D, with a scaled-down panoramic signal composited to occupy less than the top quarter of the stage STG and three different sub-scene video signals composited to occupy different amounts of the remainder of the stage STG. Figure 10B shows an alternative schematic diagram of a possible composite output or stage scene video signal, with three different sub-scene video signals adjacent to each other composited to occupy different amounts of the stage STG or composite output signal CO.

[0065] 11A and 11B show schematic diagrams of two alternative ways in which videoconferencing software may display a composite output or stage scene video signal. In FIGS. 11A and 11B, the composite output signal CO is received (e.g., via a USB port) as a single camera signal along with accompanying audio (optionally mixed and / or beamformed to emphasize the current speaker's voice) and integrated into the videoconferencing application as a single camera signal. As shown in FIG. 11A, each single camera signal is given a separate window, with the selected, active, or foreground signal, such as the composite output signal CO, reproduced as a thumbnail. In contrast, in the example shown in FIG. 11B, the selected single camera signal is given the maximum practical real estate on the display, with the selected, active, or foreground signal, such as the composite output signal CO, presented as a shaded or grayed-out thumbnail.

[0066] Sub-scene identification and synthesis 12, in step S10, new sub-scenes SS1, SS2...SSn may be generated and tracked according to the scenes, for example, as they are recognized in the panoramic video signal SC. Then, in step S30, the sub-scenes SS1, SS2...SSn may be composited according to the subject orientation, conditions, and recognition described herein. A composite output or stage scene STG,CO may then be output in step 50.

[0067] In additional detail shown in FIG. 13, and as shown in FIGS. 3A to 7B (including FIGS. 3A and 7B), in step S12, device 100 captures a wide-angle (e.g., 90-360 degree angle) scene SC with a field of view of at least 90 degrees from one or more at least partially panoramic cameras 2 or 2a...2n.

[0068] Subsequent processing for tracking and sub-scene identification may be performed on the native, undistorted, or unstitched scene SC, or on the undistorted, unstitched, or unstitched scene SC.

[0069] In step S14, new object orientations B1, B2...Bn are obtained from the wide-angle view SC using one or more of beamforming, recognition, identification, vectoring, or homing techniques.

[0070] In step S16, one or more new orientations are widened from the initial angular range (e.g., 0-5 degrees) to an angular range sufficient to span a typical human head and / or a typical human shoulder, or other default width (e.g., measured in pixels or angular range). It should be noted that the order of analysis may be reversed, e.g., a face may be detected first and then an orientation to the face may be determined. Widening may be performed in one, two, or more steps, and the two steps described herein are merely examples. Also, "widening" does not require an incremental widening process; for example, "widening" may refer to directly setting an angular range based on detection, recognition, thresholds, or values. Different methods may be used to set the angular range of a subscene. In some cases, such as when two or more faces are close to each other, the "widening" may be selected to include all of these faces, even if only one face is in the correct target orientation B1.

[0071] In step S16 (and as shown in FIGS. 5A and 5B), shoulder-width sub-scenes SS1, SS2...SSn may be set or adjusted as in step S18 according to measurements taken from the interpupillary distance or other face, head, torso, or other visible features (features, classes, colors, segments, patches, textures, trained classifiers, or other features) of the scene SC. The widths of the sub-scenes SS1, SS2...SSn may be set according to the shoulder width (alternatively according to the face width FW), or alternatively as a predetermined width related to the angular resolution of the audio microphone array 4.

[0072] Alternatively, in step S16, upper and / or lower limits on the subscene width for each or all target orientations may be set or adjusted in step S18, e.g., as peak, average, or representative shoulder width SW and face width FW, respectively. It should be noted that the notations FW and SW are used interchangeably herein as "face width" FW or "shoulder width" SW (i.e., the span of face or shoulders to be angularly captured as a subscene), and the resulting face width or shoulder width subscene SS (i.e., a block of pixels or subscene of corresponding width identified, obtained, adjusted, selected, or captured from the wide scene SC) that represents the face width FW or shoulder width SW.

[0073] In step S16, or alternatively or additionally in steps S16-S18, a first individual sub-scene of at least 20 degrees of field of view (e.g., FW1 and / or SW1) is obtained from the wide-angle scene SC at a first object orientation B1, B2...Bn. Instead of or in addition to providing at least a 20 degree field of view (e.g., FW1 and / or SW1), the first individual sub-scene FW1 and / or SW1 may be obtained from the wide-angle scene SC as a field of view spanning at least 2 to 12 times the interpupillary distance (e.g., specific to M1 or representing M1, M2...Mn), or alternatively or additionally as a field of view scaled to capture the width between the interpupillary distance (e.g., specific to M1 or representing M1, M2...Mn) and shoulder width (e.g., specific to M1 or representing M1, M2...Mn). shoulders A sub-scene capture of width SWn may record a narrower face width FWn for later reference.

[0074] If a second object orientation B1, B2...Bn is available, then in step S16, or alternatively or additionally in steps S16-S18, second individual sub-scenes (e.g., FW2 and / or SS2) are similarly obtained from the wide-angle view SC at the second object orientation, e.g., B2. If successive object orientations B3...Bn are available, then successive individual sub-scenes (e.g., FW3...n and / or SS3...n) are similarly obtained from the wide-angle view SC at the successive object orientations B3...Bn.

[0075] The first and second object orientations B1, B2 (and subsequent object orientations B3...Bn), whether obtained by stitching different camera images or from a single panoramic camera, may have a substantially common angular origin relative to the first object orientation because they are obtained from the same device 100. Optionally, one or more additional object orientations Bn from different angular origins may be obtained from another camera 5 or 7 of device 100 or from a camera on a connected device (e.g., a connected laptop, tablet, or mobile device 40 of FIG. 1A, or a connected satellite camera 7 on satellite tower 14b of FIG. 2K).

[0076] As described above, the set, obtained, or widened subscene SS, representing width FW or SW, may be adjusted in step S18, for example, (i) to be of equal or matching size to other subscenes, (ii) to be uniformly divisible or divisible (e.g., into two, three, or four segments) relative to the aspect ratio of the output image or stream signal, optionally not less than the above-mentioned lower width limit or not more than the above-mentioned upper width limit, (iii) to avoid overlap with other subscenes in nearby target orientations, and / or (iv) to be consistent in brightness, contrast, or other video characteristics with other subscenes.

[0077] In step S20 (which may include any reasonable and operable combination of steps of Modes 1, 2, or 3 from FIGS. 16-18), data and / or metadata regarding the identified object orientations B1, B2...Bn and sub-scenes FW1, FW2...FWn and / or SS1, SS2...SSn may be recorded for tracking purposes. For example, relative position from origin OR (e.g., determined by a sensor or calculation), width, height, and / or any adjusted parameters described above may be recorded.

[0078] Instead, in step S20, characteristic data, prediction data, or tracking data associated with the sub-scenes may be recorded and added, for example, to a sub-scene, orientation, or other feature tracking database in step S20. For example, sub-scenes FW1, FW2...FWn and / or SS1, SS2...SSn may be instantaneous images, image blocks, or video blocks identified within an image or video scene SC. In the case of video, prediction data may be associated with a scene or sub-scene depending on the video compression / decompression approach. It may be associated and recorded as data or metadata associated with the sub-scene, but is likely to be part of an additional new sub-scene to track.

[0079] Following recording of the tracking data or other target data, processing returns to the main routine. Composition of sub-scenes for each situation 12, processor 6 may synthesize sub-scenes SSn for each situation (e.g., for each data, flags, indicators, settings, or other action parameters recorded as tracking data or as scene data in step S20), i.e., combine first, optionally second, and optionally subsequent individual sub-scenes SSn corresponding to different widths FW1, FW2...FWn and / or SW1, SW2...SWn into a composite scene or single camera image or video signal STG or CO. In this specification, a single camera image or video signal STG, CO may refer to a single video frame or a single composite video frame representing a USB (or other peripheral bus or network) peripheral image or video signal or stream corresponding to a single USB (or other peripheral bus or network) camera.

[0080] In step S32, device 100, its circuitry, and / or its executable code may identify related subscenes SSn to be arranged as a composited combined image or video stream STG or CO. "Relevant" may be determined according to the criteria described with respect to the identification in step S14 and / or the updating and tracking in step S20. For example, one related subscene may be the subscene of the most recent speaker, and a second related subscene may be the subscene of the second-most recent speaker. These two most recent speakers may remain most relevant until a third speaker speaks and becomes more relevant. Embodiments herein accommodate three speakers within a subscene within a composite scene, each with a segment of equal width or a segment wide enough to support their head and / or shoulders. However, two speakers, four speakers, or more speakers may easily be accommodated with a wider or narrower occupancy of the composited screen width, respectively.

[0081] By selecting subscenes SSn that encapsulate faces only in height and width, up to eight speakers can be reasonably accommodated (e.g., four in the top row and four in the bottom row of the composite scene), and arrangements of four to eight speakers can be accommodated by buffering and compositing appropriate screens and / or windows (subscenes corresponding to the windows) (e.g., presenting the subscenes as an overlapping deck of cards or as foreshortened rings of view with more relevant speakers larger and closer and less relevant speakers smaller and further back). With reference to Figures 6A and 6B, scene SSn may further include whiteboard content WB whenever the system determines that the most relevant scene to display is WB (e.g., when imaged by second camera 7 as depicted in Figure 6A). The whiteboard or whiteboard scene WB may be presented prominently and occupy the majority or majority of the scene, while speakers M1, M2...Mn or SPKR may optionally be presented in picture-in-picture with the whiteboard WB content.

[0082] In step S34, the associated subscene set SS1, SS2...SSn is compared with the previously associated subscene SSn. Steps S34 and S32 may be performed in reverse order. This comparison determines whether the previously associated subscene SSn is available, should remain on the stage STG or CO, should be removed from the stage STG or CO, should be reconfigured to a smaller or larger size or perspective, or should otherwise be reconfigured to a previously composed subscene. It is determined whether a scene or stage STG or CO needs to be changed. If a new subscene SSn should be displayed, there may be too many candidate subscenes SSn for a scene change. For example, a scene change threshold may be checked in step S36 (this step may be performed before or during steps S32 and S34). For example, if the number of individual subscenes SSn becomes greater than a threshold number (e.g., 3), it may be preferable to output the entire wide-angle scene SC or a scaled-down panoramic scene SC.R (e.g., as is, or segmented and stacked to fit within the aspect ratio of the USB peripheral camera). Instead, it may be best to present a single camera scene instead of a composite scene of multiple subscenes SSn or as the composite output CO.

[0083] In step S38, device 100, its circuitry, and / or its executable code may set the subscene members SS1, SS2...SSn and the order in which they will transition and / or be composited into the composite output CO. In other words, once the candidate members of subscene complement SS1, SS2...SSn to be output as stage STG or CO and whether any rules or thresholds for scene changes have been met or exceeded have been determined, the order of the scenes SSn and the transitions in which they will be added, removed, switched, or rearranged may be determined in step S38. It should be noted that step S38 may be more or less important depending on the previous steps and the history of speakers SPKR or M1, M2...Mn. If two or three speakers M1, M2...Mn or SPKR are identified and should be displayed simultaneously as device 100 begins operation, step S38 starts with a clean slate and follows default association rules (e.g., present speakers SPKR clockwise, start with three or fewer speakers in the composite output CO). If the same three speakers M1, M2...Mn remain involved, the subscene members, order and composition may not be changed in step S38.

[0084] As described above, the identification described with respect to step S18 and the prediction / update described with respect to step S20 may result in changes to the combined output CO in steps S32-S40. In step S40, the transitions and combinations to be performed are determined.

[0085] For example, device 100 may obtain a subsequent (e.g., third, fourth, or subsequent) individual subscene SSn from wide-angle or panoramic scene SC at a subsequent target orientation. In steps S32-S38, the subsequent subscene SSn may be set to be composited or combined into a composite scene or composite output CO. Furthermore, in steps S32-S38, another subscene SSn (e.g., an earlier or less related subscene) other than the subsequent subscene may be set to be removed (by a composite transition) from the composite scene (and then composited and output as a composite scene or composite output CO that is formatted as a single-camera scene in step S50).

[0086] Additionally or alternatively, device 100 may set sub-scenes SSn to be combined or combined into or removed from the composite scene or composite output CO in steps S32-S38 according to the setting of additional criteria (e.g., utterance time, utterance frequency, audible frequency cough / sneeze / doorbell, sound amplitude, speech angle matching with face recognition) as described with reference to steps S18 and / or S20. In steps S32-S38, only subsequent sub-scenes SSn that meet the additional criteria may be set to be combined into the composite scene CO. In step S40, the transition and compositing steps to be performed are determined. The stage scene is then composited and output in step S50 as a composite output CO formatted as a single camera scene. .

[0087] Additionally or alternatively, device 100 may set subscene SSn as a protected subscene protected from removal in steps S32-S38 based on retention criteria (e.g., audio / speech duration, audio / speech frequency, time since last utterance, tagged for retention) as described with reference to steps S18 and / or S20. Removing subscene SSn other than a subsequent subscene in steps S32-S38 does not set the protected subscene to be removed from the composite scene. In step S40, the transition and compositing to be performed are determined. The composite scene is then composited and output in step S50 as a composite output CO formatted as a single-camera scene.

[0088] Additionally or alternatively, in steps S32-S38, device 100 may set sub-scene SSn enhancement operations (e.g., scaling, blinking, genie, bouncing, card sorting, ordering, cornering) as described with reference to steps S18 and / or S20 based on enhancement criteria (e.g., repeated speakers, designated presenter, most recent speaker, loudest speaker, rotating object in hand / at scene change, high-frequency scene activity in the frequency domain, hand raising). In steps S32-S38, at least one of the individual sub-scenes SSn may be set to be enhanced according to a sub-scene enhancement operation based on its respective or corresponding enhancement criteria. In step S40, the transition and compositing to be performed are determined. The composite scene is then composited and output in step S50 as a composite output CO formatted as a single-camera scene.

[0089] Additionally or alternatively, device 100 may set sub-scene participant notification or reminder actions (e.g., blinking lights for people near the sub-scene) as described with reference to steps S18 and / or S20 in steps S32-S38 based on sensors or detected criteria (e.g., too quiet, remote poke). In steps S32-S38, local reminder indicators may be set to be activated according to notification or reminder actions based on respective or corresponding detected criteria. In step S40, transitions and compositing to be performed are determined. The composite scene is then composited and output in step S50 as a composite output CO formatted as a single camera scene.

[0090] In step S40, device 100, its circuitry, and / or its executable code generate transitions and composites to smoothly change the sub-scene complement of the composite image. Following the composite output CO of the tracking data or other target data, processing returns to the main routine.

[0091] Composite Output In steps S52-S56 of FIG. 15 (optionally in reverse order), the composite scene STG or CO is formatted to be transmitted or received as a single camera scene, i.e., composited, and / or the transition is rendered or composited into a buffer, screen, or frame (where "buffer," "screen," or "frame" corresponds to the single camera view output). Device 100, its circuitry, and / or its executable code may use a compositing window or screen manager, optionally with GPU acceleration, to provide an off-screen buffer for each sub-scene, composite the buffer with surrounding and transition graphics into a single camera image representing the single camera view, and write the result to output or display memory. The compositing window or sub-screen manager circuitry may perform blending, fading, scaling, and other operations. Rotation, duplication, bending, twisting, shuffling, blurring, or other operations may be performed on the buffered windows, or drop shadows and animations such as flipping, stacking, covering, ringing, grouping, and tiling may be rendered. The compositing window manager may provide visual transitions that can be composited so that sub-scenes entering the compositing scene are added, removed, or switched with transitional effects. Sub-scenes fade in or fade out, visibly shrink in or out, and radially expand smoothly inward or outward. All scenes being composited or transitioned may be video scenes, e.g., each containing an ongoing video stream subsampled from the panoramic scene SC.

[0092] In step S52, the transition or compositing is rendered (repeatedly, incrementally, or continuously, as necessary) to a frame, buffer, or video memory (note that transition and compositing may be applied to individual frames or video streams, and may be an ongoing process through many frames of video of the entire scene STG, CO and each of the constituent sub-scenes SS1, SS2...SSn).

[0093] In step S54, device 100, its circuitry, and / or its executable code may select and transition an audio stream. Similar to the window, scene, video, or sub-scene compositing managers, the audio stream may be emphasized or de-emphasized to highlight the sub-scene being composited, particularly in the case of beam forming array 4. Similarly, synchronization of audio with the composite video scene may be performed.

[0094] In step S56, device 100, its circuitry, and / or its executable code outputs the simulation of the single camera video and audio as composite output CO. As described above, this output is at an aspect ratio and pixel count simulating a single, e.g., webcam view, of a peripheral USB device, e.g., an aspect ratio of less than 2:1 and typically less than 1.78:1, and can be used by group videoconferencing software as an external webcam input. When rendering the webcam input as a display view, the videoconferencing software treats the composite output CO as any other USB camera, and all clients interacting with host device 40 (or the version of directly connected device 100 of FIG. 1B) present the composite output CO in all main views and thumbnail views corresponding to the host device (or the version of directly connected device 100 of FIG. 1B).

[0095] Subscene compositing example 12-16, the conference camera 100 and the processor 6 may combine (at step S30) and output (at step S50) the single-camera video signals STG and CO. The processor 6, operatively connected to the ROM / RAM 8, may record (at step S12) a panoramic video signal SC having an aspect ratio of substantially 2.4:1 or greater, captured from a wide camera 2, 3, 5 having a horizontal angle of view of substantially 90 degrees or greater. In one optional version, the panoramic video signal is captured from a wide camera having an aspect ratio of substantially 8:1 or greater and a horizontal angle of view of substantially 360 degrees.

[0096] Processor 6 may (e.g., in step S14) subsample (e.g., in steps S32-S40) at least two sub-scene video signals SS1, SS2...SSn (e.g., SS2 and SS5 in Figures 8C-8E and 9C-9E) from wide camera 100 in respective target orientations B1, B2...Bn. Processor 6 may align (e.g., in steps S32-S40) two or more sub-scene video signals SS1, SS2...SSn (e.g., SS2 and SS5 in Figures 8C-8E and 9C-9E) and subsample (e.g., in steps S32-S40) The panoramic video signal SC may be composited (in a buffer, frame, or video memory) to form (in steps S52-S56) a stage scene video signal CO, STG having an aspect ratio of substantially 2:1 or less. Optionally, substantially 80% or more of the area of ​​the stage scene video signal CO, STG may be subsampled from the panoramic video signal SC to densely fill as much of the single camera video signal as possible (leading to a larger view of the participants). The processor 6 operatively connected to the USB / LAN interface 10 may output the stage scene video signal CO, STG formatted as a single camera video signal (as in steps S52-S56).

[0097] Optimally, the processor 6 may subsample additional (e.g., third, fourth, or subsequent) sub-scene video signals SS1, SS2...SS3 (e.g., SS1 in Figures 9C-9E) in their respective target orientations B1, B2...Bn from the panoramic video signal SC (and / or optionally from a buffer, frame or video memory, e.g., in the GPU 6 and / or ROM / RAM 8, and / or directly from the wide cameras 2, 3, 5). The processor may then combine two or more sub-scene video signals SS1, SS2...SS3 (e.g., SS2 and SS5 in FIGS. 9C-9E) originally combined onto the stage STG,CO with one or more additional sub-scene video signals SS1, SS2...SSn (e.g., SS1 in FIGS. 9C-9E) to form a stage scene video signal STG,CO having an aspect ratio of substantially 2:1 or less and including multiple side-by-side sub-scene video signals (e.g., two, three, four, or more sub-scene video signals SS1, SS2...SSn combined in a row or grid). The processor 6 may set or store in memory one or more target orientations or one or more additional criteria for the sub-scene video signals SS1, SS2...SSn. In this case, for example, only those additional sub-scene video signals SS1, SS2...SSn that meet the additional criteria (e.g., sufficient quality, sufficient illumination, etc.) may be transitioned to the stage scene video signal STG,CO.

[0098] Alternatively, or in addition, additional sub-scene video signals SS1, SS2...SSn may be composited into the stage scene video signal STG,CO by processor 6 by replacing one or more of the sub-scene video signals SS1, SS2...SSn that may already have been composited into the stage scene video signal STG,CO to form a stage scene video signal STG,CO that still has an aspect ratio of substantially 2:1 or less. Each sub-scene video signal SS1, SS2...SSn to be composited may be assigned a minimum width Min.1, Min.2...Min.n, and upon completion of its respective transition into the stage scene video signal STG,CO, each sub-scene video signal SS1, SS2...SSn may be composited side-by-side at substantially its minimum width Min.1, Min.2...Min.n or greater to form the stage scene video signal STG,CO.

[0099] In some cases, for example in steps S16-S18, processor 6 may increase the composite width of each sub-scene video signal SS1, SS2...SSn during the transition in an incremental manner throughout the transition until the composite width is substantially equal to or greater than its corresponding respective minimum width Min.1, Min.2...Min.n. Alternatively, or in addition, each sub-scene video signal SS1, SS2...SSn may be composited side-by-side by processor 6 at a respective width that is substantially equal to or greater than its minimum width Min.1, Min.2...Min.n, and such that the sum of all composited sub-scene video signals SS1, SS2...SSn is substantially equal to the width of the stage scene video signal or composite output STG,CO.

[0100] Alternatively, or in addition, the width of the sub-scene video signals SS1, SS2...SSn in the stage scene video signal STG,CO may be determined based on the width of one or more activities detected in one or more target orientations B1, B2...Bn corresponding to the sub-scene video signals SS1, SS2...SSn. The stage scene video signal or composite output STG,CO is synthesized by processor 6 to vary (e.g., as in steps S16-S18) according to a certain criteria (e.g., visual action, sensed action, acoustic detection of speech, etc.), whereas the width of the stage scene video signal or composite output STG,CO is kept constant.

[0101] Optionally, the processor 6 may form a stage scene video signal by combining one or more sub-scene video signals SS1, SS2...SSn (e.g., SS2 and SS5 in Figures 9C-9E) with one or more additional sub-scene video signals SS1, SS2...SSn (e.g., SS1 in Figures 9C-9E) and transitioning one or more additional sub-scene video signals SS1, SS2...SSn (e.g., SS1 in Figures 9C-9E) into the stage scene video signal STG,CO by reducing the width of one or more sub-scene video signals SS1, SS2...SSn (e.g., SS2 and SS5 in Figures 9C-9E) by an amount corresponding to the width of the one or more added or subsequent sub-scene video signals SS1, SS2...SSn (e.g., SS1 in Figures 9C-9E).

[0102] In some cases, processor 6 may assign each sub-scene video signal SS1, SS2... SSn a respective minimum width Min.1, Min.2... Min.n and may composite each sub-scene video signal SS1, SS2... SSn side-by-side at substantially equal to or greater than its corresponding respective minimum width Min.1, Min.2... Min.n to form the stage scene video signal or composite output STG,CO. When the sum of the respective minimum widths Min.1, Min.2... Min.n of two or more sub-scene video signals SS1, SS2... SSn, together with one or more additional sub-scene video signals SS1, SS2... SSn, exceeds the width of the stage scene video signal STG,CO, one or more of the two sub-scene video signals SS1, SS2... SSn may be transitioned by processor 6 to be removed from the stage scene video signal or composite output STG,CO.

[0103] In another alternative, the processor 9 may select at least one of two or more sub-scene video signals SS1, SS2...SSn to be transitioned for removal from the stage scene video signal STG,CO to correspond to the respective target orientation B1, B2...Bn in which one or more activity criteria (e.g., visual activity, detected activity, acoustic detection of speech, time since last speech, etc.) were most recently met.

[0104] In many cases, as shown in Figures 8B to 8E and Figures 9B to 9E, the processor 6 may preserve the left-to-right (clockwise when looking down) order of two or more sub-scene video signals SS1, SS2...SSn (e.g., SS2 and SS5 in Figures 9C to 9E) and one or more additional sub-scene video signals SS1, SS2...SSn (e.g., SS1 in Figures 9C to 9E) between their respective target orientations B1, B2...Bn relative to the wide cameras 2, 3, 5 when the two or more sub-scene video signals SS1, SS2...SSn are combined with at least one subsequent sub-scene video signal SS1, SS2...SSn to form a stage scene video signal or combined output STG,CO.

[0105] Alternatively, or in addition, processor 6 may select each of the target orientations B1, B2... Bn from the panoramic video signal SC depending on one or more selection criteria (e.g., visual activity, sensed activity, acoustic detection of speech, time since last speech, etc.) detected in each of the target orientations B1, B2... Bn relative to wide cameras 2, 3, 5. After one or more selection criteria are no longer true, processor 6 may transition the corresponding sub-scene video signal SS1, SS2... SSn to be removed from the stage scene video signal or composite output STG, CO. The selection criteria may include the existence of an activity criterion met in each of the target orientations B1, B2... Bn. Processor 9 may count the time since one or more activity criteria were met in each of the target orientations B1, B2... Bn. A predetermined period of time after one or more activity criteria are met in the elephant orientation B1, B2...Bn, the processor 6 may transition the respective sub-scene signal SS1, SS2...SSn to be removed from the stage scene video signal STG.

[0106] With respect to the downsized panoramic video signal SC.R shown in Figures 8A-8C, 9A-9C, 10A, 1B, 11A, 11B, and 22, processor 6 may subsample the downsized panoramic video signal SC.R from the panoramic video signal SC to produce a downsized panoramic video signal SC.R with an aspect ratio of substantially 8:1 or greater. Processor 6 may then combine two or more sub-scene video signals (e.g., SS2 and SS5 in Figures 8C-8E and 9C-9E) with the downsized panoramic video signal SC.R to form a stage scene video signal STG,CO having an aspect ratio of substantially 2:1 or less, including a plurality of side-by-side sub-scene video signals (e.g., SS2 and SS5 in Figures 8C-8E, SS1, SS2, and SS5 in Figures 9C-9E) and the panoramic video signal SC.R.

[0107] In this case, the processor 6 may combine two or more sub-scene video signals (e.g., SS2 and SS5 in Figures 8C-8E, SS1, SS2 and SS5 in Figures 9C-9E) together with a scaled-down panoramic video signal SC.R to form a stage scene video signal having an aspect ratio of substantially 2:1 or less, including a plurality of side-by-side sub-scene video signals (e.g., SS2 and SS5 in Figures 8C-8E, SS1, SS2 and SS5 in Figures 9C-9E) and a panoramic video signal SC.R that is taller than the plurality of side-by-side sub-scene video signals, and the panoramic video signal is no more than 1 / 5 of the area of ​​the stage scene video signal or composite output STG or CO and extends substantially across the width of the stage scene video signal or composite output STG or CO.

[0108] 24, processor 6 may subsample text video signal TD1 from, or be provided with, a text document (e.g., from a text editor, word processor, spreadsheet, presentation, or other document that renders text), and processor 6 may then transition the text video signal TD1 or a rendered or scaled down version thereof, TD1.R, into stage scene video signal STG,CO by replacing at least one of the two or more sub-scene video signals with text video signal TD1 or an equivalent, TD1.R.

[0109] Optionally, processor 6 may set one or more of the two sub-scene video signals as protected sub-scene video signals SS1, SS2...SSn that are protected from transitioning based on one or more retention criteria (e.g., visual motion, detected motion, acoustic detection of speech, time since last speech, etc.) In this case, processor 6 may transition one or more additional sub-scene video signals SS1, SS2...SSn into the stage scene video signal by replacing at least one of the two or more sub-scene video signals SS1, SS2...SSn, but in particular by transitioning out sub-scene video signals SS1, SS2...SSn other than the protected sub-scene.

[0110] Alternatively, processor 6 may set a sub-scene emphasis operation (e.g., blinking, highlighting, outlining, icon overlay, etc.) based on one or more emphasis criteria (e.g., visual activity, detected activity, acoustic detection of speech, time since last speech, etc.), in which case one or more sub-scene video signals are emphasized according to the sub-scene emphasis operation and based on the corresponding emphasis criteria.

[0111] In a further variation, the processor 6 may use a sensor-detected criterion (e.g., sound waves, vibrations, electric currents detected by a sensor such as an RF element, a passive infrared element, or a distance recognition element). The sub-scene participant notification action may be set based on the detected criteria (magnetic radiation, heat, UV radiation, radio, microwave, electrical properties, or depth / range detection). The processor 6 may activate one or more local reminder indicators according to the notification action based on the corresponding detected criteria.

[0112] Examples of target directions For example, the target orientations may be those orientations corresponding to one or more audio signals or detections, such as, for example, speaking participants M1, M2... Mn, as angle-detected, vectored, or identified by the microphone array 4, e.g., by beamforming, localization, or comparative received signal strength, or comparative time-of-flight using at least two microphones. Thresholding or frequency domain analysis may be used to determine whether the audio signals are strong enough or clear enough, and filtering may be performed using at least three microphones to discard mismatched pairs, multipath, and / or redundancies. Three microphones have the advantage of forming three pairs for comparison.

[0113] As another example, alternatively or additionally, the target orientation may be those orientations in which motion is detected, angle-recognized, vectorized, or identified within a scene by features, images, patterns, classes, and / or motion detection circuitry or executable code capable of scanning images or motion video or RGBD from camera 2.

[0114] As another example, instead or in addition, the subject orientation may be those orientations at which facial structures are detected, angled, vectorized, or otherwise identified within a scene by a face detection circuit or executable code capable of scanning images or motion video or RGBD signals from camera 2. Skeletal structures may also be detected in this manner.

[0115] As another example, instead or in addition, the object orientations may be those orientations in which substantially continuous structures of color, texture, and / or pattern are detected, angle recognized, vectorized, or identified in a scene by edge detection, corner detection, blob detection or segmentation, extrema detection, and / or feature detection circuitry or executable code capable of scanning images or motion video or RGBD signals from camera 2. The recognition may refer to previously recorded, learned, or trained image patches, colors, textures, or patterns.

[0116] As another example, alternatively or additionally, the target orientations may be those orientations in a scene that are detected, angle-recognized, vectorized, or otherwise identified as being different from a known environment by difference and / or change detection circuitry or executable code capable of scanning images or motion video or RGBD signals from camera 2. For example, device 100 may maintain one or more visual maps of an empty conference room in which the device is located and detect that a sufficiently obstructing entity, such as a person, is obstructing a known feature or area in the map.

[0117] As another example, alternatively or additionally, the target orientation may be the orientation in which regular shapes such as "whiteboard" shapes, door shapes, or rectangles, including the shape of a chair back, are identified, angle recognized, vectorized, or otherwise identified by features, images, patterns, classes, and / or motion detection circuitry or executable code capable of scanning images or motion video or RGBD from camera 2.

[0118] As another example, alternatively or additionally, the target orientation may include active or passive acoustic emitters or transducers, and / or passive or active optical or visual fiducial markers, and / or RFID or other electromagnetically detectable. Reference objects or features recognizable as artificial landmarks may be those oriented as placed by a person using device 100, which are angle recognized, vectorized, or identified by one or more of the techniques described above.

[0119] If an initial or new target orientation cannot be obtained in this way (e.g., because none of the participants M1, M2...Mn have yet spoken), a default view can be set to be output as a single-camera scene instead of a composite scene. For example, as one default view, the entire panoramic scene (e.g., with an H:V horizontal-vertical ratio of 2:1 to 10:1) can be fragmented and arranged in the output single-camera ratio (e.g., an H:V aspect ratio or horizontal-vertical ratio of 1.25:1 to 2.4:1 or 2.5:1, typically in a landscape orientation, although a corresponding "reverse" portrait orientation ratio is also possible). As another example, a default view in front of the target orientation can be first obtained, and a "window" corresponding to the output scene ratio can be tracked at a fixed rate across the entire scene SC, e.g., as a simulation of a slowly panning camera. As another example, the default view can consist of "headshots" of each conference attendee M1, M2...Mn (with an additional 5-20% width in the margins), with the margins adjusted to optimize the available display area.

[0120] Aspect Ratio Examples While embodiments and aspects of the invention may be useful with any angle range or aspect ratio, advantages are optionally greater when sub-scenes are formed from cameras providing panoramic video signals having an aspect ratio of substantially 2.4:1 or greater (aspect ratio refers to either frame or pixel dimensions) and are composited into a multi-participant stage video signal having an overall aspect ratio of substantially 2:1 or less (e.g., 16:9, 16:10, or 4:3), as seen on most laptop or television displays (typically 1.78:1 or less), and, optionally, when the stage video signal sub-scenes fill more than 80% of the area of ​​the composited overall frame, and / or when the stage video signal sub-scenes and any additionally composited thumbnails formed from the panoramic video signal fill more than 90% of the area of ​​the composited overall frame. In this way, each shown participant fills the screen as nearly as possible as practical.

[0121] The corresponding ratio between the vertical and horizontal angles of view can be determined from α=2 arctangent as the ratio (d / 2f), where d is the vertical or horizontal dimension of the sensor and f is the effective focal length of the lens. Different wide-angle cameras for conferences may have a 90-, 120-, or 180-degree field of view from a single lens, but each camera may output a 1080p image (e.g., a 1920x1080 image) with an aspect ratio of 1.78:1, or a much wider image with an aspect ratio of 3.5:1, or other aspect ratios. When observing a conference scene, a smaller aspect ratio (e.g., 2:1 or less) combined with a 120- or 180-degree wide camera may show more of the ceiling, walls, or tables than might be desired. Thus, although the aspect ratio of the scene or panoramic video signal SC and the field of view FOV of the camera 100 may be independent, it is optionally advantageous in this embodiment to match wider cameras 100 (90 degrees or more) with video signals of wider aspect ratios (e.g., 2.4:1 or more), and optionally, the widest cameras (e.g., 360 degree panoramic views) with the widest aspect ratios (e.g., 8:1 or more).

[0122] Subscene or Orientation Tracking Example 1A and 1B, as shown in Figures 12-18, and particularly Figures 16-18, may involve tracking sub-scenes FW, SS at target orientations B1, B2, ... Bn within a wide video signal SC. As shown in Figure 16, an acoustic sensor or microphone array 4 (with optional beamforming circuitry) and wide cameras 2, 3, 5 The processor 6 operatively connected to monitors a substantially common angular range, optionally or preferably substantially 90 degrees or greater, in step S202.

[0123] In steps S204 and S206, processor 6 may execute code or include or be operatively connected to circuitry that identifies a first object orientation B1, B2... Bn along localization (e.g., a measurement representing a position in Cartesian or polar coordinates, or in a direction) of one or both of acoustic recognition (e.g., frequency, pattern, or other audio recognition) and visual recognition (e.g., motion detection, face detection, bone structure detection, color blob segmentation or detection) within the angular range of wide cameras 2, 3, 5. As in step S10 and as in steps S12 and S14, a sub-scene video signal SS is subsampled from wide cameras 2, 3, 5 along the object orientation B1, B2... Bn identified in step S14 (e.g., newly sampled from the image sensors of wide cameras 2, 3, 5 or subsampled from the panoramic scene SC captured in step S12). The width of the sub-scene video signal SS (e.g., minimum width Min.1, Min.2...Min.n, or sub-scene display width DWid.1, DWid.2...DWid.n) may be set by the processor 6 in step S210 according to signal characteristics of one or both of the acoustic and visual / visual recognition. The signal characteristics may represent various acoustic or visual recognition quality or confidence levels. As used herein, "acoustic recognition" may include any recognition based on sound waves or vibrations (e.g., meeting a measurement threshold, matching a descriptor, etc.), including waveform frequency analysis such as Doppler analysis, while "visual recognition" may include any recognition corresponding to electromagnetic radiation (e.g., meeting a measurement threshold, matching a descriptor, etc.), such as heat or UV radiation, radio or microwave, electrical property recognition, or depth / range detected by sensors such as RF elements, passive infrared elements, or distance recognition elements.

[0124] For example, the target orientations B1, B2...Bn identified in step S14 can be determined by combining such acoustic and visual recognition in different orders, some of which are shown in Figures 16-18 as Modes 1, 2, or 3 (which can be reasonably and logically combined with each other). In one order, the acoustic recognition orientation is recorded first (although this order can be repeated and / or changed), as in step S220 of Figure 18. Optionally, such orientations B1, B2...Bn can be orientations of a certain angle, a certain angle with a tolerance, or an approximate range or angular range (such as orientation B5 of Figure 7A). As shown in steps S228-S232 of Figure 18, the recorded acoustic recognition orientation can be refined (narrowed or reevaluated) based on visual recognition (e.g., face recognition) if a sufficiently reliable visual recognition is substantially within a threshold angular range of the recorded acoustic recognition. In the same mode, or combined with another mode, for example as in step S218 of FIG. 17, any acoustic recognition not associated with a visual recognition may remain a candidate target orientation B1, B2...Bn.

[0125] Optionally, as in step S210 of FIG. 16 , the signal characteristics represent a confidence level for one or both of the acoustic and visual recognition. "Confidence level" need not meet a formal probabilistic definition but may refer to any comparative measure that establishes a degree of reliability (e.g., exceeding a threshold amplitude, signal quality, signal-to-noise ratio or equivalent, or success criteria). Alternatively, or in addition, as in step S210 of FIG. 16 , the signal characteristics may represent the width of a recognized feature within one or both of the acoustic recognition (e.g., the angular range in which a sound may originate) or visual recognition (e.g., interpupillary distance, face width, body width). For example, the signal characteristics may correspond to the approximate width of a recognized human face (e.g., determined by visual recognition) along the target orientations B1, B2...Bn. The width of the first sub-scene video signals SS1, SS2...SSn may be set according to the signal characteristics for visual recognition.

[0126] In some cases, such as step S228 of FIG. 18 , if the width cannot be set according to the signal characteristics of visual recognition (e.g., if the width cannot be reliably set if the width-defining feature cannot be recognized), a predetermined width may be set along the localization of detected acoustic recognition within an angular range, such as step S230 of FIG. 18 . For example, if a face cannot be recognized by image analysis along target orientations B1, B2...Bn evaluated as having an acoustic signal indicative of human speech, such as steps S228 and S232 of FIG. 18 , a default width (e.g., a subscene having a width equivalent to 1 / 10 to 1 / 4 of the width of the entire scene SC) may be maintained or set along the acoustic orientation to define subscene SS, such as step S230. For example, FIG. 7A illustrates a scenario of attendees and speakers in which attendee M5's face is facing the direction of attendee M4 and M5 is speaking. In this case, the acoustic microphone array 4 of the conference camera 100 may be able to localize the speaker M5 along the target orientation B5 (here, the target orientation B5 is depicted as an orientation range rather than a vector), but image analysis of the panoramic scene SC of the wide camera 2, 3, 5 video signal may not be able to resolve faces or other visual recognition. In such a case, a default width Min.5 may be set as the minimum width for initially defining, constraining, or rendering the sub-scene SS5 along the target orientation B5.

[0127] In another embodiment, the target orientations B1, B2... Bn may be identified as directed toward an acoustic recognition detected within the angular range of the conference camera 100. In this case, the processor 6 may optionally identify a visual recognition proximate to the acoustic recognition (e.g., within, overlapping, or adjacent to the target orientations B1, B2... Bn, e.g., within 5-20 degrees of the arc of the target orientations B1, B2... Bn) as in step S209 of FIG. 16 . In this case, the width of the first sub-scene video signals SS1, SS2... SSn may be set according to the signal characteristics of the visual recognition that was (or is) proximate to or otherwise matched with the acoustic recognition. This may occur, for example, when the target orientations B1, B2... Bn are first identified by the acoustic microphone array 4 and then validated or confirmed with a sufficiently close or otherwise matching facial recognition using video images from the wide camera 100.

[0128] 17 and 16, a system including conference or wide camera 100 may use latent visual or acoustic recognition to create a spatial map, as in step S218 of FIG. 17, and then rely on this spatial map to verify the validity of subsequent associated, matching, adjacent, or "snap" recognitions using the same or different or other recognition approaches, as in step S209 of FIG. 16. For example, in some cases, the entire panoramic scene SC may be too large to effectively scan frame-by-frame for purposes such as facial recognition. In this case, because people do not move significantly in a conference situation using camera 100, especially after they have sat down at their desks for the conference, only a portion of the entire panoramic scene SC may be scanned, for example, per video frame.

[0129] For example, as in step S212 of FIG. 17, to track sub-scenes SS1, SS2...SSn at object orientations B1, B2...Bn within the wide video signal, processor 6 may scan a sub-sampling window through the motion video signal SC corresponding to a wide camera 100 field of view of substantially 90 degrees or greater. Processor 6 or circuitry associated therewith may identify candidate object orientations B1, B2...Bn within the sub-sampling window by substantially satisfying a threshold for defining suitable signal quality for the candidate object orientations B1, B2...Bn, as in step S214 of FIG. 17. Each object orientation B1, B2...Bn may correspond to a localization of visual perception detected within the sub-sampling window, as in step S216 of FIG. 17. As in step S218 of FIG. 17, The candidate orientations B1, B2... Bn may then be recorded in a spatial map (e.g., a memory or database structure that keeps track of the position, location, and / or direction of the candidate orientations). In this manner, for example, facial recognition or other visual recognition (e.g., motion) may be stored in the spatial map even if no acoustic detection has yet occurred in that orientation. The angular range of the wide camera 100 may then be monitored by the processor 6 using an acoustic sensor or microphone array 4 for acoustic recognition (which may be used to verify the validity of the candidate target orientations B1, B2... Bn).

[0130] For example, referring to FIG. 7A , the processor 6 of the conference camera 100 may scan different subsampled windows across the panoramic scene SC for visual recognition (e.g., faces, colors, movements, etc.). Depending on lighting, movements, facial orientation, etc., potential object orientations corresponding to face, movement, or similar detections of attendees M1...M5 in FIG. 7 may be stored in a spatial map. However, in the scenario shown in FIG. 7A , a potential object orientation on the side of attendee Map.1 may not be later validated by acoustic signals if it corresponds to an attendee who is not speaking (and this attendee may not be captured in the subscene at all, but only in the panoramic scene). When attendees M1...M5 speak or begin to speak, potential object orientations including or on the side of these attendees may be validated and recorded as object orientations B1, B2...B5.

[0131] Optionally, as in step S209 of FIG. 16 , when acoustic recognition is detected in close proximity (substantially adjacent, next to, or within an arc of + / - 5 to 20 degrees) to a candidate orientation recorded in the spatial map, processor 6 may snap target orientations B1, B2...Bn to substantially correspond to the candidate orientation. Step S209 of FIG. 16 indicates matching the target orientation with the corresponding spatial map, where "matching" may include correlating, replacing, or modifying the target orientation value. For example, face or action recognition within a window and / or panoramic scene SC may have better resolution than an acoustic array or microphone array 4, but because detection is less frequent or reliable, the detected target orientations B1, B2...Bn resulting from acoustic recognition may be modified, recorded, or otherwise corrected or adjusted according to visual recognition. In this case, instead of subsampling the sub-scene video signals SS1, SS2...SSn along the apparent object orientations B1, B2...Bn obtained from acoustic recognition, the processor 6 may subsample the sub-scene video signals along the object orientations B1, B2...Bn following a snap motion, e.g., from the wide camera 100 and / or the panoramic scene SC after the acoustic object orientations B1, B2...Bn have been corrected using previously mapped visual recognition. In this case, as in step S210 of Figure 16, the width of the sub-scene video signals SS may be set according to the detected face width or action width, or alternatively according to signal characteristics of the acoustic recognition (e.g., default width, resolution of array 4, confidence level, width of features recognized in one or both of the acoustic or visual recognition, approximate width of the recognized human face along the object orientation). If the sub-scene SS width is not set according to the signal characteristics of visual recognition, such as face width or movement range, as in step S210 of FIG. 16 or step S230 of FIG. 18, a predetermined width (such as the default width Min. 5 as in FIG. 7A) may be set according to acoustic recognition.

[0132] In the example of FIG. 18, the conference camera 100 and processor 6 may track sub-scenes in target orientations B1, B2...Bn by recording motion video signals corresponding to a field of view FOV of the wide camera 100 of substantially 90 degrees or more. The processor may monitor an angular range corresponding to the field of view FOV of the wide camera 100 using the acoustic sensor array 4 for acoustic recognition in step S220, and once a range of acoustic recognition is detected in step S222, may identify target orientations B1, B2...Bn oriented toward the detected acoustic recognition within that angular range in step S224. The processor 6 and associated circuitry may then, in step S226, determine a target orientation B1, B2...Bn (e.g., the same as the range of target orientation B5 in FIG. 7A) that corresponds to the angular range of the wide camera 100. 18 , processor 6 may position a sub-sampling window within the motion video signal of the panoramic scene SC according to a corresponding range of object orientations B1, B2, ... Bn (such as B1, B2, ... Bn). If visual recognition is detected within that range, processor 6 may then localize the detected visual recognition within the sub-sampling window, as in step S228. Processor 6 may then sub-sample a sub-scene video signal SS captured from wide camera 100 (either directly from camera 100 or from panoramic scene recording SC) optionally substantially centered on the visual recognition. Processor 6 may then set the width of the sub-scene video signal SS according to the signal characteristics of the visual recognition, as in step S232. If visual recognition is not possible, preferred, not detected, or not selected, as in step S228 of FIG. 18 , processor 6 may maintain or select an acoustic minimum width, as in step S230 of FIG. 18 .

[0133] 16-18, the conference camera 100 and processor 6 may track a sub-scene at object orientations B1, B2... Bn within a wide video signal, such as a panoramic scene SC, by monitoring an angular range using the acoustic sensor array 4 and wide cameras 2, 3, 5 observing a field of view of substantially 90 degrees or more, as in step S212 of FIG. 17. The processor 6 may identify multiple object orientations B1, B2... Bn, each oriented toward a localization (acoustic, visual, or sensor-based, as in step S216) within the angular range, and may maintain a spatial map of recorded characteristics corresponding to the object orientations B1, B2... Bn as the object orientations B1, B2... Bn, corresponding recognitions, corresponding localizations, or data representing same are sequentially stored as in step S218 of FIG. 17. Thereafter, for example as in step S210 of FIG. 16, the processor 6 may subsample the sub-scene video signals SS1, SS2...SSn from the wide camera 100 substantially along at least one target orientation B1, B2...Bn, and set the width of the sub-scene video signals SS1, SS2...SSn according to the recorded characteristics corresponding to the at least one target orientation B1, B2...Bn.

[0134] Predictive Tracking Example The above description of structures, devices, methods, and techniques for identifying new object orientations describes various detection, recognition, triggering, or other causes for identifying such new object orientations. The following description discusses updating, tracking, or predicting changes in the orientation, direction, location, pose, width, or other characteristics of object orientations and subscenes, but this updating, tracking, and prediction may also apply to the above description. The description of methods for identifying new object orientations and updating or predicting changes in orientation or subscenes is relevant in that tracking or prediction facilitates reacquisition of the object orientation or subscene. The methods and techniques described herein can be used to scan, identify, update, track, record, or reacquire the orientation and / or subscene in steps S20, S32, S54, or S56, or vice versa.

[0135] For example, predictive video data may be recorded for each sub-scene, such as data encoded according to or associated with predicted HEVC, H.264, MPEG-4, other MPEG I slices, P slices, and B slices (or frames, or macroblocks); other intra-frame and inter-frame pictures, macroblocks, or slices; H.264 or other SI frame / slice, SP frame / slice (switching P), and / or multi-frame motion prediction; VP9 or VP10 superblocks, blocks, macroblocks or superframes, intra-frame and inter-frame prediction, component prediction, motion compensation, motion vector prediction, and / or segmentation.

[0136] For example, motion vectors derived from audio motion on a microphone array, or direct or pixel-based methods (e.g., block matching, phase correlation, frequency domain correlation, pixel recursion, optical flow) and / or indirect or feature-based methods (e.g., sub-microphones). Other prediction or tracking data as described above independent of the video standard or motion compensation SPI may be recorded, such as motion vectors obtained from feature detection (e.g., corner detection using statistical functions such as RANSAC applied over a scene or scene region).

[0137] Additionally or alternatively, the subscene-by-subscene updating or tracking may record, identify, or score relevant indicators or data or information representative thereof, such as obtained audio parameters, for example, amplitude, frequency of vocalizations, length of vocalizations, associated attendees M, M... M (two subscenes with reciprocal traffic), moderator or moderator attendee M.Lead (a subscene with periodic brief audio interjections), recognized signal phase (e.g., applause, "keep the camera on me," and other expressions and speech recognition). These parameters or indicators may be recorded independently of the tracking step or at different times during the tracking step. The subscene-by-subscene tracking may also record, identify, or score erroneous or irrelevant indicators, such as, for example, audio representative of coughing or sneezing; regular or periodic motion or video representative of machinery, wind, or flashing; transient motion or motion at a frequency high enough to be transient;

[0138] Additionally or alternatively, per-subscene updating or tracking may record, identify, or score indicators or data or information representative thereof for setting a subscene and / or protecting the subscene from removal, for example, based on retention criteria (e.g., voice / utterance time, voice / utterance frequency, time since last utterance, tagged for retention). In subsequent processing for compositing, removing a subscene other than a new or subsequent subscene will not remove the protected subscene from the composite scene. In other words, the protected subscene will have a low priority for removal from the composite scene.

[0139] Additionally or alternatively, the sub-scene-by-sub-scene updating or tracking may record, identify, or score indicators or representative data or information (e.g., time of speech, frequency of speech, audible frequency cough / sneeze / doorbell, sound amplitude, speech angle match with facial recognition) to set additional criteria, and only subsequent sub-scenes that meet the additional criteria are combined into a composite scene during compilation.

[0140] Additionally or alternatively, the per-subscene updating or tracking may record, identify, or score indicators or data or information representative thereof for setting subscene emphasis actions, such as audio, CGI, image, video, or compositing effects, based on emphasis criteria (e.g., repeat speaker, designated presenter, most recent speaker, loudest speaker, motion detection of rotating object in hand / scene change, high frequency scene activity in the frequency domain, hand raise motion, or skeletal recognition) (e.g., scaling one subscene larger, blinking or pulsing the border of one subscene, inserting a new subscene with a genie effect (increasing from small to large), highlighting or inserting a subscene with a bounce effect, arranging one or more subscenes with a card sorting or shuffle effect, ordering subscenes with an overlap effect, cornering a subscene with a "folded" graphic corner appearance). In the compilation process, at least one of the individual subscenes is highlighted according to a subscene emphasis action based on its or corresponding emphasis criteria.

[0141] Additionally or alternatively, per-subscene updates or tracking may record, identify, or score indicators or data or information representative thereof for setting sub-scene participant notification or reminder actions based on sensors or detected criteria (e.g., too quiet, remote pokes from social media) (e.g., turning on lights on devices 100 for attendees M1, M2...Mn, optionally lights on the same side as the sub-scene). During the compilation process or other processes, local reminder indicators are activated according to notification or reminder actions based on respective or corresponding detected criteria.

[0142] Additionally or alternatively, the updating or tracking for each sub-scene may, for example, record, identify or score indicators or data or information representative thereof for predicting or setting a change vector for each angular sector FW1, FW2...FWn or SW1, SW2...SWn based on a change in speed or direction of the recorded characteristic of each recognition or localization (e.g., a color blob, face, voice as described herein with respect to step S14 or S20), and / or for updating the direction of each angular sector FW1, FW2...FWn or SW1, SW2...SWn based on said prediction or setting.

[0143] Additionally or alternatively, the updating or tracking for each subscene may, for example, record, identify, or score indicators or data or information representative thereof for predicting or setting a search area for recapturing or reacquiring a lost recognition or localization based on the most recent location of a recorded characteristic (e.g., color blob, face, voice) of each recognition or localization, and / or for updating the orientation of each angular sector based on said prediction or setting. The recorded characteristic may be at least one color blob, segmentation, or blob object representing skin and / or clothing.

[0144] Additionally or alternatively, the sub-scene-by-sub-scene updating or tracking may maintain a Cartesian map or, in particular or optionally, a polar map (e.g., based on orientations B1, B2...Bn or angles from an origin OR in the scene SC and angular ranges such as sub-scenes SS1, SS2...SSn corresponding to angular sectors FW, SW in the scene SC) of the recorded characteristics, each of which has at least one parameter representing the orientation B1, B2...Bn of the recorded characteristic.

[0145] Thus, alternatively or additionally, an embodiment of device 100, its circuitry, and / or executable code stored and executing within ROM / RAM 8 and / or CPU / GPU 6 may track target sub-scenes SS1, SS2...SSn corresponding to widths FW and / or SW within wide-angle scene SC by monitoring a target angular range (e.g., the horizontal extent of cameras 2n, 3n, 5, or 7 forming scene SC, or a subset thereof) with acoustic sensor array 4 and optical sensor arrays 2, 3, 5, and / or 7. Device 100, its circuitry, and / or its executable code may scan target angular range SC for recognition criteria (e.g., sounds, faces), e.g., as described herein with respect to step S14 (identifying new target orientations) and / or step S20 (tracking and characteristic information for orientations / sub-scenes) of FIG. 8 . Device 100, its circuitry, and / or its executable code may identify a first object orientation B1 based on a first recognition (e.g., detection, identification, triggering, or other cause) and localization (e.g., angle, vector, pose, or location) by acoustic sensor array 4 and at least one of optical sensor arrays 2, 3, 5, and / or 7. Device 100, its circuitry, and / or its executable code may identify a second object orientation B2 (and optionally third and subsequent object orientations B3...Bn) based on a second recognition and localization (and optionally third and subsequent recognition and localization) by acoustic sensor array 4 and at least one of optical sensor arrays 2, 3, 5, and / or 7.

[0146] The device 100, its circuitry, and / or its executable code may be configured to generate angular sub-scenes (e.g., an initial small angular range or a face-based sub-scene FW) containing respective target orientations B1, B2, ... Bn, based on at least one recognition criterion (e.g., a set or reset angular range). The respective angular sector (e.g., FW, SW, or others) for each target orientation B1, B2...Bn may be set by expanding, widening, setting, or resetting until a threshold (e.g., a width threshold as described with reference to steps S16-S18 of Figure 13) based on (e.g., the degree span is wider than, twice the interpupillary distance, or wider; the set or reset angular span is wider than head-to-wall contrast, distance, edge, difference, or motion transition) is met.

[0147] Device 100, its circuitry, and / or its executable code may update or track (these terms are used interchangeably herein) the direction or orientation B1, B2... Bn of the respective angular sectors FW1, FW2... FWn and / or SW1, SW2... SWn based on changes in the direction or orientation B1, B2... Bn within each recognition and / or localization or of recorded characteristics (e.g., color blobs, faces, sounds) representative of each recognition and / or localization. Optionally, as described herein, device 100, its circuitry, and / or its executable code may update or track the respective angular sectors FW1, FW2... FWn and / or SW1, SW2... SWn to follow angular changes in the first, second, and / or third, and / or subsequent object orientations B1, B2... Bn.

[0148] Example of composite output (video conference) In FIGS. 8A-8D, 10A-10B, and 19-24, the "composite output CO," i.e., the combined or composited sub-scene as a composited rendered / composite camera view, is shown with a callout to both the main view of remote display RD1 (representing the scene received from conference room local display LD) and network interface 10 or 10a, indicating that the videoconferencing client on conference room (local) display LD "transparently" treats the video signal received from USB peripheral device 100 as a single camera view and conveys the composite output CO to the remote client or remote displays RD1 and RD2. Note that all thumbnail views may also show the composite output CO. In general, FIGS. 19, 20, and 22 correspond to the arrangement of attendees shown in FIGS. 3A-5B, with an additional attendee joining in FIG. 21, seated in the empty seat shown in FIGS. 3A-5B.

[0149] During an exemplary transition, the scaled-down panoramic video signal SC.R (occupying approximately 25% of the vertical screen) may show a "zoomed-in" portion of the panoramic scene video signal SC (e.g., as shown in Figures 9A-9E). The zoom level may be determined by the number of pixels contained in this approximately 25%. As people / objects M1, M2...Mn become relevant, the corresponding sub-scenes SS1, SS2...SSn transition into the stage scene STG or composite output CO (e.g., by compositing a sliding video panel), maintaining their clockwise or left-to-right position between participants M1, M2...Mn. Simultaneously, the processor, using the GPU 6 memory or ROM / RAM 8, may slowly scroll the scaled-down panoramic video signal SC.R left or right to display the current target orientation B1, B2...Bn in the center of the screen. The current target orientation may be highlighted. As new relevant sub-scenes SS1, SS2...SSn are identified, the scaled-down panoramic video signal SC.R may be rotated or panned such that the most recent sub-scene SS1, SS2...SSn is highlighted and positioned at the center of the scaled-down panoramic video signal SC.R. With this configuration, the scaled-down panoramic video signal SC.R is continuously re-rendered and effectively panned throughout the conference to show the relevant portion of the room.

[0150] As shown in Figure 19, in a typical video conference display, each participant's display shows a master view and multiple thumbnail views, each substantially determined by the output signal of a webcam. The master view is typically one of the remote participants, and the thumbnail views represent the other participants. In a video conference or chat system, Depending on the system, the master view may be selected to show the active speaker in the audience, or may be switched to another audience, often by thumbnail selection, in some cases including a local scene. In some systems, the local scene thumbnail remains in the overall display at all times so that each audience member can position themselves relative to the camera to present a useful scene (an example of this is shown in Figure 19).

[0151] As shown in Figure 19, an embodiment of the present invention provides a composite stage view of multiple attendees instead of a single camera scene. For example, in Figure 19, attendees M1, M2, and M3 (represented by icons M1, M2, and M3) 3 Potential target orientations B1, B2, and B3 to are available to the conference camera 100. As described herein, because there are three possible attendees M1, M2, and M3 that are localized or otherwise identified, and one SPKR is speaking, the stage STG (equivalent to the composite output CO) may initially be populated with a default number (two in this case) of relevant subscenes, including a subscene of the active speaker SPKR, which in FIG. 19 is attendee M2.

[0152] 19 shows the displays of three participants: a local display LD, such as a personal computer connected to the conference camera 100 and to the Internet INET; a first personal computer (“PC”) or tablet display remote display RD1 of a first remote attendee A.hex; and a second PC or tablet display RD2 of a second remote attendee A.diamond. As expected in the context of a video conference, the local display LD most prominently shows the remote speaker selected by the operator of the local display PC or by the video conference software (A.hex in FIG. 19), while the two remote displays RD1, RD2 show views selected by the remote operator or software (e.g., the view of the active speaker, the composite view CO of the conference camera 100).

[0153] 19, the local display LD typically shows a master view with the most recently selected remote attendee (e.g., A.hex, the attendee working at the PC or laptop with remote display RD1) and a thumbnail column (including a composited stage view from the local conference camera 100) in which essentially all attendees are represented. Each of the remote displays RD1 and RD2, in contrast, shows a master view with a composited stage view CO, STG (since speaker SPKR is currently speaking), and the thumbnail column again contains views of the remaining attendees.

[0154] FIG. 19 assumes that attendee M3 has already spoken or was previously selected as the default occupant of stage STG and is already occupying the most relevant subscene (e.g., the most recently relevant subscene). As shown in FIG. 19, subscene SS1 corresponding to speaker M2 (icon M2 and silhouette M2 with an open mouth on remote display 2) is composited into a single camera view with a slide transition (represented by a block arrow). A preferred slide transition starts with zero or negligible width, slides the middle, i.e., target orientations B1, B2, Bn of corresponding subscenes SS1, SS2, ... SSn onto the stage, then grows the width of the composited corresponding subscenes SS1, SS2, ... SSn until it reaches at least the minimum width, and may continue to grow the width of the composited corresponding subscenes SS1, SS2, ... SSn until the entire stage is filled. The composite (middle transition) and composited scene are provided as camera views to the videoconferencing client on the meeting room (local) display LD, so the composite and composited scene may be presented substantially simultaneously (i.e., presented as the current view) in the main view and thumbnail view of the local client display LD and the two remote client displays RD1, RD2.

[0155] In Figure 20, after Figure 19, attendee M1 becomes the most recent and / or most relevant speaker (e.g., the previous situation was that of Figure 19, where attendee M2 was the most recent and / or most relevant speaker). Subscenes SS3 and SS2 for attendees M3 and M2 remain relevant according to tracking and identification criteria and can be reconfigured to smaller widths as needed (by scaling or cropping, optionally limited by a width limit of 2-12 times the interpupillary distance and other methods as described herein). Subscene SS2 is similarly composited to a compatible size and then composed onto stage STG with a sliding transition (again represented by a block arrow). As described herein with respect to Figures 9, 10A-10B, and 11A-11B, the new speaker SPKR is attendee M1 to the right (clockwise in a top-down view) of the orientation of already displayed attendee M2, so subscene SS1 may optionally be transitioned onto the stage to preserve the left-right or left-to-right order (M3, M2, M1), in this case transitioning from the right.

[0156] In FIG. 21, new attendee M4, who arrived in the room after FIG. 20, becomes the most recent and relevant speaker. Subscenes SS2 and SS1 for speakers M2 and M1 remain relevant according to tracking and identification criteria and remain composited at a "3:1" width. The subscene corresponding to speaker M3 "ages out" and is no longer as relevant as the most recent speaker (although many other priorities and associations are described herein). Subscene SS4 corresponding to speaker M4 is composited to a compatible size and then composited onto the camera output with a flip transition (again represented by a block arrow), while subscene SS3 is flipped out for removal. This may be a slide transition or an alternate transition. Alternatively, although not shown, because new speaker SPKR is attendee M4 to the left (clockwise in a top-down view) of the orientation of already-displayed attendees M2 and M1, subscene SS4 may optionally be transitioned onto the stage to preserve the lateral view or left-to-right order (M4, M2, M1), in this case, transitioning from the left. In this case, subscenes SS2 and SS1 may each transition one place to the right, and subscene M3 may exit to the right of the stage (as in a sliding transition away).

[0157] 19-21 illustrate exemplary local and remote video conferencing modes, on a mobile device by way of example, in which a composited, tracked, and / or displayed composite scene is received and displayed as a single camera scene, as described herein and referenced in the context of the previous paragraph.

[0158] While the overall information is similar, Fig. 22 presents a form for displaying a video conference that is a variation of the form of Fig. 19. In particular, whereas in Fig. 19 the thumbnail views do not overlap the master view and thumbnail views that match the master view are maintained in the thumbnail column, in the form of Fig. 22 the thumbnails overlap the master view (e.g., are composited so as to be superimposed on the master view) and the current master view is not highlighted in the thumbnail column (e.g., by dimming, etc.).

[0159] FIG. 23 shows a fourth client corresponding to a separate high-resolution, close-up, or similar camera 7 connecting its client to the videoconference group via network interface 10b, while the composite output CO and its transitions are presented on a conference room (local) display LD via network interface 10a. 19 to 22 show variations of the above.

[0160] FIG. 24 shows a variation of FIGS. 19-22 in which a client reviewing code or a document with a text review window connects to the conference camera 100 via a local wireless connection (although in some variations, the client reviewing the code or document may connect via the Internet from a remote station). In one example, a first device or client (PC or tablet) runs a video conferencing or chat client showing attendees in a panoramic view, and a second client or device (PC or tablet) runs a code or document review client and provides it to the conference camera 100 as a video signal in the same form as a webcam. The conference camera 100 composites the document window / video signal of the code or document review client as a full-frame subscene SSn onto the stage STG or CO, and optionally also composites a local panoramic scene containing the conference attendees, e.g., higher than the stage STG or CO. In this way, instead of individual attendee subscenes, the text shown in the video signal is available to all attendees, but attendees may still be identified by referring to the panoramic view SC. Although not shown, the conference camera 100 device may alternatively create, instantiate, or run a second videoconferencing client to host the document view. Alternatively, a high-resolution, close-up, or simply separate camera 7 connects its own client to the videoconferencing group via network interface 10b, while the composite output CO and its transitions are presented to the conference room (local) display via network interface 10a.

[0161] In at least one embodiment, conferees M1, M2... Mn may always be shown in the stage scene video signal or composite output STG,CO. Based on at least face width detection, processor 6 may crop faces as face-only sub-scenes SS1, SS2... SSn and align them along the top or bottom of the stage scene video signal or composite output STG,CO, as shown in FIG. 25 . In this case, it may be desirable for a participant using a device such as remote device RD1 to click or (if a touchscreen) touch the cropped face-only sub-scenes SS1, SS2... SSn to communicate with local display LD to create a stage scene video signal STG centered on that person. In one exemplary solution, using a configuration directly connected to the Internet INET similar to FIG. 1B , conference camera 100 may create or instantiate an appropriate number of virtual videoconference clients and / or assign a virtual camera to each.

[0162] FIG. 26 illustrates some of the iconography and symbols used throughout FIGS. 1-26. In particular, arrows extending from the center of a camera lens may correspond to target orientations B1, B2, ... Bn, regardless of whether the arrows are labeled as such in the various figures. Dashed lines extending from a camera lens at open "V" angles may correspond to the lens's field of view, regardless of whether the dashed lines are labeled as such in the various figures. A schematic "stick figure" depiction of a person with an oval head and a rectangular or trapezoidal body may correspond to a conference participant, regardless of whether the schematic person is labeled as such in the various figures. A depiction of the schematic person's open mouth may depict the current speaker SPKR, regardless of whether the schematic person with the open mouth is labeled as such in the various figures. A thick arrow extending from left to right, right to left, top to bottom, or spiraling may indicate an ongoing transition or a combination of transitions, regardless of whether the arrows are labeled as such in the various figures.

[0163] In this disclosure, "wide-angle camera" and "wide scene" depend on the field of view and distance from the subject, and are suitable for capturing two different people in a meeting who are not shoulder-to-shoulder. This includes any camera with a sufficiently wide field of view.

[0164] A "field of view" is the horizontal field of view of a camera unless a vertical field of view is specified. As used herein, a "scene" refers to an image (still or video) of a scene captured by a camera. Generally, but with exceptions, a panoramic "scene" SC is one of the largest images or video streams or signals handled by a system, regardless of whether the signal is captured by a single camera or stitched from multiple cameras. The most commonly referred to scene "SC" herein includes a scene SC that is a panoramic scene SC captured by a camera coupled to a fisheye lens, a camera coupled to panoramic optics, or an equiangular distribution of overlapping cameras. The panoramic optics may provide the panoramic scene substantially directly to the camera; in the case of a fisheye lens, the panoramic scene SC may be a periphery or horizontal band of the fisheye view separated and dewarped into a long, high-aspect-ratio rectangular image; in the case of overlapping cameras, the panoramic scene may be stitched and cropped (and possibly dewarped) from the individual overlapping views. A "sub-scene" refers to a sub-portion of a scene, e.g., a contiguous, usually rectangular, block of pixels smaller than the entire scene. A panoramic scene may be cropped to less than 360 degrees and still be referred to as the entire scene SC within which the sub-scene is addressed.

[0165] As used herein, "aspect ratio" is described as an H:V horizontal:vertical ratio, with "larger" aspect ratios increasing the horizontal ratio (wider and shorter) relative to the vertical. Aspect ratios greater than 1:1 (e.g., 1.1:1, 2:1, 10:1) are considered "landscape format," and for purposes of this disclosure, aspects of 1:1 or less are considered "portrait format" (e.g., 1:1.1, 1:2, 1:3). A "single camera" video signal is formatted as a video signal corresponding to a single camera, e.g., UVC, also known as the "USB Device Class Definition for Video Devices" 1.1 or 1.5 by the USB Implementers Forum, each of which is incorporated herein by reference in its entirety (i.e., see http: / / www.usb.org / developers / docs / devclass_docs / USB_Video_Class_1_5.zip USB_Video_Class_1_1_090711.zip at the same URL). Any of the signals described in UVC can be a "single camera video signal" regardless of whether the signal is transported, carried, transmitted, or tunneled over USB.

[0166] "Display" means any direct display screen or projection display. "Camera" means a digital imaging device, which may be a CCD or CMOS camera, a thermal imaging camera, or an RGBD depth or time-of-flight camera. The camera may be formed by two or more stitched camera views and / or a wide aspect, panoramic, wide-angle, fisheye, or catadioptric perspective virtual camera.

[0167] A "participant" is a person, device, or place connected to a group video conferencing session and displaying a view from a webcam; in most cases, an "attendee" is not only a participant, but is also in the same room as the conference camera 100. A "speaker" is an attendee who is speaking or has spoken recently enough for the conference camera 100 or an associated remote server to identify the speaker, although in some descriptions it may also be a participant who is speaking or has spoken recently enough for the video conferencing client or an associated remote server to identify the speaker.

[0168] "Compositing" generally refers to digital compositing as known in the art, i.e., the digital assembly of multiple video signals (and / or images or other media objects) to create a final video signal, which includes alpha compositing and blur compositing. Compositing includes techniques such as rendering, anti-aliasing, node-based compositing, keyframing, layer-based compositing, nesting compositing or composites, and deep image compositing (using color, opacity, and depth using deep data, whether function-based or sample-based). Compositing is an ongoing process involving the movement and / or animation of sub-scenes, each containing a video stream; for example, various frames, windows, and sub-scenes within an overall stage scene may each display different ongoing video streams as they move, transition, blend, or otherwise composite as the overall stage scene. Compositing, as used herein, may use a compositing window manager with one or more off-screen buffers for one or more windows, or a stacking window manager. Any off-screen buffer or display memory contents may be double- or triple-buffered or otherwise buffered. Compositing may further include operations on either or both the buffered window or the display memory window, such as applying 2D and 3D animation effects, blending, fading, scaling, zooming, rotating, duplicating, bending, twisting, shuffling, blurring, drop shadows, glows, previews, and adding animations. Compositing may also include applying these to vector-oriented or pixel- or voxel-oriented graphical elements. Compositing may include rendering pop-up previews upon touch, mouseover, hover, or click; window switching by rearranging several windows against a background to allow selection by touch, mouseover, hover, or click; and flip, cover, ring, exposure, and the like. As described herein, various visual transitions may be used on the stage, such as fading, sliding, growing or shrinking, and combinations thereof. As used herein, "transition" includes the necessary compositing steps.

[0169] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in a user terminal.

[0170] All of the above processes may be embodied in, and fully automated via, software code modules executed by one or more general-purpose or special-purpose computers or processors. The code modules may be stored on any type of computer-readable medium or other computer storage device or collection of storage devices. Alternatively, some or all of the methods may be embodied in specialized computer hardware.

[0171] All of the methods and tasks described herein may be performed and fully automated by a computer system. The computer system may, in some cases, include multiple individual computers or computing devices (e.g., physical servers, workstations, storage arrays, etc.) communicating and interoperating over a network to perform the functions described above. Each such computing device typically includes a processor. A computer system may include a processor (or multiple processors or a circuit or collection of circuits, e.g., a module) that executes program instructions, or a module stored in memory or other non-transitory computer-readable storage medium. While various functions disclosed herein may be embodied in such program instructions, some or all of the disclosed functionality may instead be realized in application-specific circuitry (e.g., an ASIC or FPGA) of a computer system. When a computer system includes multiple computing devices, these devices may, but need not, be co-located. The results of the disclosed methods and tasks may be persistently stored by converting physical storage devices, such as solid-state memory chips and / or magnetic disks, to a different state.

Claims

1. A method for tracking a sub-scene in a target orientation within a video signal, comprising: monitoring the angular range with an acoustic sensor array including a plurality of microphones configured to capture the angular range and a camera observing a field of view of 90 degrees or more; identifying a first target orientation along a localization of at least one of detected acoustic or visual perceptions within the angular range; sub-sampling a first sub-scene video signal from the camera along the first object orientation; setting a width of the first sub-scene video signal based on one or more of frequency detection or frequency pattern of the acoustic recognition within the angular range, or motion detection, face detection, bone structure detection, or color blob segmentation of the visual recognition within the angular range; determining the width of the first sub-scene video signal based on a width of a feature recognized in at least one of the acoustic recognition or the visual recognition.

2. The method of claim 1 , comprising determining the width of the first sub-scene video signal based on an approximate width of a recognized human face along the first object orientation.

3. The method of claim 1 , comprising setting a predetermined width along the localization of the detected acoustic perception within the angular range.

4. A method for tracking a sub-scene in a target orientation within a video signal, comprising: monitoring the angular range with an acoustic sensor array including a plurality of microphones configured to capture the angular range and a camera observing a field of view of 90 degrees or more; identifying a first target orientation along a localization of at least one of detected acoustic or visual perceptions within the angular range; sub-sampling a first sub-scene video signal from the camera along the first object orientation; setting a width of the first sub-scene video signal based on one or more of frequency detection or frequency pattern of the acoustic recognition within the angular range, or motion detection, face detection, bone structure detection, or color blob segmentation of the visual recognition within the angular range; determining the first target orientation based on the visual recognition; and and setting the width of the first sub-scene video signal based on the visual perception.

5. A method for tracking a sub-scene in a target orientation within a video signal, comprising: monitoring the angular range with an acoustic sensor array including a plurality of microphones configured to capture the angular range and a camera observing a field of view of 90 degrees or more; identifying a first target orientation along a localization of at least one of detected acoustic or visual perceptions within the angular range; sub-sampling a first sub-scene video signal from the camera along the first object orientation; setting a width of the first sub-scene video signal based on one or more of frequency detection or frequency pattern of the acoustic recognition within the angular range, or motion detection, face detection, bone structure detection, or color blob segmentation of the visual recognition within the angular range; identifying the visual recognition as being proximate to the acoustic recognition; and setting the width of the first sub-scene video signal based on the visual perception proximate to the audio perception.

Citation Information

Patent Citations

  • Imaging apparatus

    JP2003092726A

  • Method and system for automatic detection and tracking of multiple individuals using multiple cues

    JP2003216951A

  • Video conference system with minute creation support function

    JP2005341015A

  • Photographing device and communication conference system

    JP2007124140A

  • Conference image reproducing apparatus and conference image reproducing method

    JP2009182980A