Composite and scale angle-separated sub-scenes
By using a wide camera to capture panoramic video signals and subsample multiple subsceive video signals in corresponding position of interest, the processor synthesizes these subsceive video signals to form stage scene video signals, solving the problem of difficult to automatically synthesize and display subsceived angle separation in wide scenes in video conferences, and improving the experience quality of remote participants.
Patent Information
- Application Number
- CN202111304450.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-04-01
- Filing Date
- 2016-04-01
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2036-04-01
AI Technical Summary
The prior art is difficult to automatically and seamlessly synthesize, track and display angle-separated subscenes and subscenes of interest within wide scenes in video conferencing, making it difficult for remote participants to understand all participants in the conference room.
By capturing panoramic video signals using a wide camera and subsampling multiple subsceive video signals at the corresponding direction of interest, the processor combines these subsceive video signals side by side to form a stage scene video signal, formatting them into a single camera video signal.
It realizes automatic synthesis and display of angle separation in wide scenes in video conferencing, improving the sense of participation and experience quality of remote participants.
Smart Images

Figure CN114422738B_ABST
Abstract
Description
[0001] This application is a divisional application. The original application was a patent application submitted to the China Patent Office on November 30, 2017 (the international application date is April 1, 2016), with application number 201680031904.8, and the name of the invention is “Synthesizing and scaling angle-separated sub-scenes”.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the benefit under 35 USC §119(e) of U.S. Provisional Patent Application Serial No. 62 / 141,822, filed on April 1, 2015, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0004] Various aspects relate to apparatus and methods for image capture and emphasis. Background Art
[0005] Multi-party teleconferencing, video chats, and teleconferences are often conducted with multiple participants in a conference room connected to at least one remote party.
[0006] In the case of person-to-person mode of video conferencing software, only one local camera is often available with a limited horizontal field of view (e.g., 70 degrees). Whether the single camera is located in front of one participant or at the head of a table facing all participants, it is difficult for the remote party to understand the audio, body language, and non-verbal cues given by those participants in the conference room who are away from the single camera or at an acute angle to the camera (e.g., seeing the side of a person instead of the face).
[0007] In the case of multiplayer mode of video conferencing software, the availability of cameras of two or more mobile devices (laptops, tablets or mobile phones) located in the same conference room adds a few different problems. The more conference room participants logged into the meeting, the greater the audio feedback and crosstalk becomes. The camera viewing angle may be as far away from the participants or skewed as in the case of a single camera. Local participants may tend to engage with other participants through their mobile devices despite being in the same room (thus suffering the same weaknesses as the remote party in terms of body language and non-verbal cues).
[0008] There are no known commercial or experimental technologies for synthesizing, tracking and / or displaying angularly separated sub-scenes and / or sub-scenes of interest within a wide scene (e.g., a wide scene of two or more conference participants) in a manner that makes setup very easy for participants in the same room or that makes the experience automatic and seamless from the perspective of remote participants. Summary of the invention
[0009] In one aspect of this embodiment, the process of outputting a densely synthesized single camera signal can record a panoramic video signal having an aspect ratio of substantially 2.4:1 or greater, which is captured from a wide camera having a horizontal field of view angle of substantially 90 degrees or greater. At least two sub-scene video signals can be sub-sampled from the wide camera at corresponding bearings of interest. Two or more sub-scene video signals can be synthesized side by side to form a stage scene video signal having an aspect ratio of substantially 2:1 or less. Optionally, an area of more than 80% of the stage scene video signal is sub-sampled from the panoramic video signal. The stage scene video signal can be formatted as a single camera video signal. Optionally, the panoramic video signal has an aspect ratio of substantially 8:1 or greater, which is captured from a wide camera having a horizontal field of view angle of substantially 360 degrees.
[0010] In a related aspect of this embodiment, the conference camera is configured to output a densely synthesized single camera signal. The imaging element or wide camera of the conference camera can be configured to capture and / or record a panoramic video signal having an aspect ratio of substantially 2.4:1 or greater, and the wide camera has a horizontal field of view angle of substantially 90 degrees or greater. A processor operably connected to the imaging element or the wide camera can be configured to subsample two or more sub-scene video signals from the wide camera at corresponding azimuths of interest. The processor can be configured to synthesize the two or more sub-scene video signals as side-by-side video signals into a memory (e.g., a buffer and / or video memory) to form a stage scene video signal having an aspect ratio of substantially 2:1 or less. The processor can be configured to synthesize the sub-scene video signals into a memory (e.g., a buffer and / or video memory) so that an area of more than 80% of the stage scene video signal is subsampled from the panoramic video signal. The processor can also be configured to format the stage scene video signal into a single camera video signal, such as for transmission via USB.
[0011] In any of the above aspects, the processor may be configured to perform sub-sampling of the additional sub-scene video signals at corresponding azimuths of interest from the panoramic video signal, and to synthesize the two or more sub-scene video signals with the one or more additional sub-scene video signals to form a stage scene video signal including a plurality of side-by-side sub-scene video signals and having an aspect ratio of substantially 2:1 or less. Optionally, synthesizing the two or more sub-scene video signals with the one or more additional sub-scene video signals to form the stage scene video signal includes transitioning the one or more additional sub-scene video signals into the stage scene video signal by replacing at least one of the two or more sub-scene video signals to form the stage scene video signal having an aspect ratio of substantially 2:1 or less.
[0012] Further optionally, each sub-scene video signal may be assigned a minimum width, and after completing each respective transfer to the stage scene video signal, each sub-scene video signal may be synthesized side by side at substantially no less than its minimum width to form the stage scene video signal. Alternatively or additionally, the synthesized width of each respective sub-scene video signal that is transferred may increase throughout the entire transfer process until the synthesized width is substantially equal to or greater than the corresponding respective minimum width. Further alternatively or additionally, the sub-scene video signals may be synthesized side by side at substantially no less than their minimum width, and each sub-scene video signal may be synthesized at a respective width at which the sum of all synthesized sub-scene video signals is substantially equal to the width of the stage scene video signal.
[0013] In some cases, the width of the sub-scene video signal within the stage scene video signal can be synthesized to change according to the activity criteria detected at one or more locations of interest corresponding to the sub-scene video signal, while the width of the stage scene video signal remains constant. In other cases, the two or more sub-scene video signals are synthesized with the one or more additional sub-scene video signals to form the stage scene video signal, including transferring the one or more additional sub-scene video signals into the stage scene video signal by reducing the width of at least one of the two or more sub-scene video signals by an amount corresponding to the width of the one or more additional sub-scene video signals.
[0014] Further optionally, each sub-scene video signal may be assigned a corresponding minimum width, and each sub-scene video signal may be synthesized side by side with a width substantially not less than the corresponding corresponding minimum width to form the stage scene video signal. When the sum of the corresponding minimum widths of two or more sub-scene video signals and one or more additional sub-scene video signals exceeds the width of the stage scene video signal, at least one of the two or more sub-scene video signals may be shifted to be removed from the stage scene video signal. Optionally, the sub-scene video signal shifted to be removed from the stage scene video signal corresponds to the corresponding location of interest that least recently satisfies the activity criterion.
[0015] In any of the above aspects, when two or more sub-scene video signals are synthesized with one or more additional sub-scene video signals to form a stage scene video signal, the left-to-right order of the two or more sub-scene video signals and the one or more additional sub-scene video signals in their corresponding directions of interest relative to the wide camera can be maintained.
[0016] Furthermore, in any of the above aspects, each respective position of interest from the panoramic video signal may be selected based on a selection criterion detected at the respective position of interest relative to the wide camera. After the selection criterion is no longer true, the corresponding sub-scene video signal may be transferred to be removed from the stage scene video signal. Alternatively or in addition, the selection criterion may include the presence of an activity criterion satisfied at the respective position of interest. In this case, the processor may calculate a time from when the activity criterion was satisfied at the respective position of interest. A predetermined period of time after the activity criterion is satisfied at the respective position of interest, the corresponding sub-scene signal may be transferred to be removed from the stage scene video signal.
[0017] In a further variation of the above aspect, the processor may perform subsampling of a reduced panoramic video signal having an aspect ratio of substantially 8:1 or greater from the panoramic video signal, and synthesize two or more sub-scene video signals with the reduced panoramic video signal to form a stage scene video signal including a plurality of side-by-side sub-scene video signals and the panoramic video signal and having an aspect ratio of substantially 2:1 or less. Optionally, the two or more sub-scene video signals may be synthesized with the reduced panoramic video signal to form a stage scene video signal including a plurality of side-by-side sub-scene video signals and a panoramic video signal above the plurality of side-by-side sub-scene video signals and having an aspect ratio of substantially 2:1 or less, the panoramic video signal not exceeding 1 / 5 of the area of the stage scene video signal and extending substantially across the width of the stage scene video signal.
[0018] In a further variation of the above aspect, the processor or associated processor may sub-sample a text video signal from a text document and transfer the text video signal to a stage scene video signal by replacing at least one of the two or more sub-scene video signals with the text video signal.
[0019] Optionally, the processor may set at least one of the two or more sub-scene video signals as a protected sub-scene video signal that is protected from being transferred based on a retention criterion. In this case, the processor may transfer one or more additional sub-scene video signals to the stage scene video signal by replacing at least one of the two or more sub-scene video signals and / or by transferring sub-scene video signals other than the protected sub-scene.
[0020] In some cases, the processor may alternatively or additionally set a subscene emphasis operation based on an emphasis criterion, wherein at least one of the two or more subscene video signals is emphasized according to the subscene emphasis operation based on the corresponding emphasis criterion. Optionally, the processor may set a subscene participant notification operation based on a sensing criterion from a sensor, wherein a local reminder indicium (e.g., light, flash, or sound) is activated according to the notification operation based on the corresponding sensing criterion.
[0021] In one aspect of this embodiment, a process for tracking subscenes at locations of interest within a wide video signal may include monitoring an angular range with an acoustic sensor array and a wide camera that observes a field of view that is substantially 90 degrees or greater. A first location of interest may be identified along the location of at least one of acoustic recognition and visual recognition detected within the angular range. A first subscene video signal may be subsampled from the wide camera along the first location of interest. A width of the first subscene video signal may be set based on a signal characteristic of at least one of the acoustic recognition and the visual recognition.
[0022] In a related aspect of this embodiment, the conference camera may be configured to output a video signal including a sub-scene subsampled and scaled from a wide-angle scene, and track the sub-scenes and / or locations of interest within the wide video signal. The conference camera and / or its processor may be configured to monitor the angular range with an acoustic sensor array and a wide camera that observes a field of view of substantially 90 degrees or greater. The processor may be configured to identify a first location of interest along the location of at least one of acoustic recognition and visual recognition detected within the angular range. The processor may be further configured to subsample the first sub-scene video signal from the wide camera to a memory (buffer or video) along the first location of interest. The processor may also be configured to set the width of the first sub-scene video signal based on the signal characteristics of at least one of the acoustic recognition and visual recognition.
[0023] In any of the above aspects, the signal characteristic may represent a confidence level of either or both of the acoustic recognition or the visual recognition. Optionally, the signal characteristic may represent a width of a feature identified within either or both of the acoustic recognition or the visual recognition. Further optionally, the signal characteristic may correspond to an approximate width of a face identified along the first orientation of interest.
[0024] Alternatively or in addition, when the width is not set according to the signal characteristics of the visual recognition, the predetermined width can be set along the positioning of the acoustic recognition detected within the angular range. Further optionally, the first orientation of interest can be determined by visual recognition, and the width of the first subscene video signal is then set according to the signal characteristics of the visual recognition. Also optionally, the first orientation of interest can be identified as pointing to the acoustic recognition detected within the angular range. In this case, the processor can identify a visual recognition that is close to the acoustic recognition, and the width of the first subscene video signal can then be set according to the signal characteristics of the visual recognition that is close to the acoustic recognition.
[0025] In another aspect of this embodiment, the processor may be configured to perform a process of tracking a subscene at a location of interest within a wide video signal, including scanning a subsampling window across a moving video signal corresponding to a wide camera field of view of substantially 90 degrees or greater. The processor may be configured to identify candidate locations within the subsampling window, each location of interest corresponding to a location of visual recognition detected within the subsampling window. The processor may then record the candidate locations in a spatial map, and may use an acoustic sensor array for acoustic recognition to monitor an angular range corresponding to the wide camera field of view.
[0026] Optionally, when an acoustic recognition is detected close to a candidate position recorded in the spatial map, the processor may further snap a first position of interest to substantially correspond to a candidate position, and may subsample a first subscene video signal from a wide camera along the first position of interest. Optionally, the processor may also be configured to set a width of the first subscene video signal based on a signal characteristic of the acoustic recognition. Further optionally, the signal characteristic may represent a confidence level of the acoustic recognition; or may represent a width of a feature recognized within either or both of the acoustic recognition or the visual recognition. The signal characteristic may alternatively or additionally correspond to an approximate width of a face recognized along the first position of interest. Optionally, when the width is not set based on a signal characteristic of the visual recognition, a predetermined width may be set along the position of the acoustic recognition detected within an angular range.
[0027] In another aspect of this embodiment, the processor can be configured to track subscenes at locations of interest, including by recording motion video signals corresponding to a wide camera field of view that is substantially 90 degrees or greater. The processor can be configured to monitor an angular range corresponding to the wide camera field of view using an acoustic sensor array for acoustic recognition, and identify a first location of interest directed to an acoustic recognition detected within the angular range. A subsampling window can be positioned in the motion video signal based on the first location of interest, and visual recognition can be detected within the subsampling window. Optionally, the processor can be configured to subsample a first subscene video signal captured from a wide camera that is substantially centered around the visual recognition, and set the width of the first subscene video signal based on the signal characteristics of the visual recognition.
[0028] In another aspect of this embodiment, the processor can be configured to track subscenes at locations of interest within a wide video signal, including monitoring the angular range using an acoustic sensor array and a wide camera that observes a field of view of substantially 90 degrees or greater. A plurality of locations of interest can be identified, each location pointing to a location within the angular range. The processor can be configured to maintain a spatial map having recording characteristics corresponding to the locations of interest, and to subsample the subscene video signal from the wide camera substantially along one or more locations of interest. The width of the subscene video signal can be set according to the recording characteristics corresponding to at least one location of interest.
[0029] In another aspect of this embodiment, the processor may be configured to perform a process of tracking subscenes at locations of interest within a wide video signal, including monitoring an angular range using an acoustic sensor array and a wide camera that observes a field of view of substantially 90 degrees or greater, and identifying multiple locations of interest that all point to a location within the angular range. The subscene video signal may be sampled from the wide camera substantially along at least one location of interest, and the width of the subscene video signal may be set by extending the subscene video signal until a threshold based on at least one identification criterion is met. Optionally, a change vector for each location of interest may be predicted based on a change in one of a speed and a direction of a recorded characteristic corresponding to the location, and the position of the location of interest may be updated based on the prediction. Optionally, a search area for location may be predicted based on a recent position of a recorded characteristic corresponding to the location, and the location of the location may be updated based on the prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1A and 1B is a schematic block diagram of an embodiment of a device suitable for synthesizing, tracking and / or displaying angularly separated sub-scenes and / or sub-scenes of interest within a wide scene acquired by the device 100.
[0031] Figures 2A to 2K is used for Figure 1A and 1B The device 100 is a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement suitable for capturing a wide scene and / or a panoramic scene.
[0032] Figure 3A and 3B A top view and a conference camera panoramic image signal of a conference camera use case showing three participants are respectively shown.
[0033] Figure 4A and 4B A top view and a conference camera panoramic image signal are respectively shown for a conference camera use case showing a conference table, showing three participants, and including a depiction of the recognition of a face width setting or sub-scene.
[0034] Figure 5A and 5B A top view and a conference camera panoramic image signal are respectively shown for a conference camera use case showing a conference table, showing three participants, and including a depiction of the identification of a shoulder width setting or sub-scene.
[0035] Fig. 6A and 6B A top view and a conference camera panoramic image signal are shown, respectively, of a conference camera use case showing a conference table, showing three participants and a whiteboard, and including a depiction of the recognition of a wider sub-scene.
[0036] Fig. 7A and 7B A top view and a conference camera panoramic image signal are respectively shown for a conference camera use case showing a conference table with ten seats, showing five participants, and including a depiction of the identification of the visual minimum width and orientation and the auditory minimum width and orientation.
[0037] Fig. 8A A schematic diagram showing the extraction of a panoramic video signal and a sub-scene video signal to be synthesized into a stage scene video signal, a minimum width, and a conference camera video signal is shown.
[0038] Figure 8B A schematic diagram showing a panoramic video signal and a sub-scene video signal to be synthesized into a stage scene video signal is shown, and Figures 8C to 8E Three possible composite output or stage scene video signals are shown.
[0039] Fig.9A A schematic diagram showing a substitute panoramic video signal to be synthesized into a stage scene video signal, extraction of a substitute sub-scene video signal, a minimum width, and a conference camera video signal is shown.
[0040] Fig. 9B Schematic diagrams of alternative panoramic video signals and alternative sub-scene video signals to be synthesized into a stage scene video signal are shown, and 9C to 9E show three possible alternative synthesized outputs or stage scene video signals.
[0041] Fig.9F A schematic diagram of a panoramic video signal adjusted so that the conference table image is arranged in a more natural, less awkward view is shown.
[0042] Fig. 10A and 10B A schematic diagram showing possible composite output or stage scene video signals.
[0043] Fig.11A and 11B Schematic diagram showing two alternative ways in which video conferencing software can display composite output or stage scene video signals.
[0044] Fig.12 A flow chart comprising steps for compositing one or more stage scene video signals is shown.
[0045] Fig.13 A detailed flow chart including steps for synthetically creating sub-scenes (sub-scene video signals) based on locations of interest is shown.
[0046] Fig.14 A detailed flow chart including steps for compositing sub-scenes into a stage scene video signal is shown.
[0047] Fig.15 A detailed flow chart including steps for outputting a composite stage scene video signal as a single camera signal is shown.
[0048] Fig.16 A detailed flow chart of a first mode including performing steps for positioning and / or positioning of interest and / or setting the width of a sub-scene is shown.
[0049] Fig.17 A detailed flow chart of the second mode including performing steps for locating and / or positioning of interest and / or setting the width of a sub-scene is shown.
[0050] Fig.18 A detailed flow chart of a third mode including performing steps for positioning and / or a position of interest and / or setting a width of a sub-scene is shown.
[0051] Figure 19-21 shows that it basically corresponds to Figures 3A-5BOperation of an embodiment includes a conference camera attached to a local PC having a video conferencing client receiving a single camera signal, the PC in turn being connected to the Internet, and two remote PCs also receiving the single camera signal within a video conferencing display.
[0052] Fig. 22 Shows Figure 19-21 A variation of the system in which the video conferencing client uses overlapping video views rather than discrete adjacent views.
[0053] Fig.23 shows that it basically corresponds to Figure 6A-6B of Figure 19-21 A variation of the system that includes a high-resolution camera view for the whiteboard.
[0054] Fig.24 Shows Figure 19-21 A variation of the system including a high-resolution text document view (e.g., a text editor, word processing, presentation, or spreadsheet).
[0055] Fig.25 Is used with Figure 1B A similar configuration is shown in Figure 1, in which a schematic diagram of an arrangement of video conferencing clients is instantiated for each sub-scenario.
[0056] Fig.26 is a schematic diagram of some exemplary diagrams and symbols used throughout FIGS. 1-26 . DETAILED DESCRIPTION
[0057] Conference Camera
[0058] Figure 1A and 1B is a schematic block diagram of an embodiment of a device suitable for synthesizing, tracking and / or displaying angularly separated sub-scenes and / or sub-scenes of interest within a wide scene captured by the device, the conference camera 100 .
[0059] Figure 1AA device is shown that is constructed to communicate as a conference camera 100 or conference "webcam", for example as a USB peripheral connected to a USB host or hub of a connected laptop, tablet or mobile device 40; and provides a single video image of an aspect ratio, pixel count and proportion commonly used by existing video chat or video conferencing software (e.g., "Google Hangouts", "Skype" or "Facetime"). Device 100 includes a "wide camera" 2, 3 or 5, for example, a camera capable of capturing more than one conference participant and pointed to observe the conference participants or participants M1, M2...Mn. Camera 2, 3 or 5 may include one digital imager or lens, or 2 or more digital imagers or lenses (e.g., stitched in software or otherwise). It should be noted that depending on the location of device 100 within the conference, the field of view of wide camera 2, 3 or 5 may not exceed 70 degrees. However, in one or more embodiments, the wide camera 2, 3, 5 may be used in the center of the meeting, and in this case, the wide camera may have a horizontal field of view of approximately 90 degrees or greater than 140 degrees (not necessarily continuous), or up to 360 degrees.
[0060] In a large conference room (e.g., one designed to fit more than 8 people), it may be useful to have multiple wide-angle camera devices record a wide field of view (e.g., approximately 90 degrees or greater) and cooperatively stitch together a very wide scene to capture the most pleasing angle; for example, a wide-angle camera at the far end of a long (10-20 foot) table may result in an unsatisfactory distant view of speaker SPKR, but having multiple cameras distributed over the table (e.g., 1 camera per 5 seats) may produce at least one satisfactory or pleasing view. Cameras 2, 3, 5 may image or record a panoramic scene (e.g., with an aspect ratio of 2.4:1 to 10:1, such as a ratio of H:V horizontal to vertical) and / or make that signal available via a USB connection.
[0061] As about Figures 2A-2K As discussed, the height of the wide cameras 2, 3, 5 from the bottom of the conference camera 100 is preferably greater than 8 inches, so that the cameras 2, 3, 5 can be taller than a typical laptop screen at a meeting, thereby having an unobstructed and / or approximately eye-level view for the participants M1, M2...Mn. The microphone array 4 includes at least two microphones and can obtain an interesting orientation of nearby sounds or speeches through beamforming, relative time of flight, positioning, or received signal strength differences as known in the art. The microphone array 4 can include multiple microphone pairs that are pointed to cover an angular range at least substantially the same as the field of view of the wide camera 2.
[0062] The microphone array 4 may optionally be arranged at a height greater than 8 inches along with the wide cameras 2, 3, 5, also so that there is a direct "line of sight" between the array 4 and the participants M1, M2...Mn when they are speaking, without being blocked by a typical laptop screen. A CPU and / or GPU (and associated circuitry, such as camera circuitry) 6 for processing computational and graphical events is connected to each of the wide cameras 2, 3, 5 and the microphone array 4. ROM and RAM 8 are connected to the CPU and GPU 6 for storing and receiving executable code. A network interface and stack 10 is provided for USB, Ethernet and / or WiFi connection to the CPU 6. One or more serial buses interconnect these electronic components, and they are powered by DC, AC or battery power.
[0063] The camera circuitry of cameras 2, 3, 5 may output processed or rendered images or video streams as a single camera image signal, video signal or stream having an "H:V" horizontal to vertical ratio or aspect ratio of from 1.25:1 to 2.4:1 or 2.5:1 in the horizontal direction (e.g., including 4:3, 16:10, 16:9 ratios), and / or, as noted, with appropriate lenses and / or stitching circuitry, output a panoramic image or video stream as a single camera image signal of substantially 2.4:1 or greater. Figure 1A The conference camera 100 can be connected typically as a USB peripheral device to a laptop, tablet or mobile device 40 (having a display, network interface, computing processor, memory, camera and microphone parts, interconnected by at least one bus), on which multi-party teleconferencing, video conferencing or video chat software resides, and can be connected to a remote client 50 via the Internet 60 for teleconferencing.
[0064] Figure 1B yes Figure 1A A variant of Figure 1A The device 100 and the teleconferencing device 40 are both connected. The camera circuit output as a single camera image signal, video signal or video stream is directly available to the CPU, GPU, associated circuits and memory 5, 6, and the teleconferencing software is resident on behalf of the CPU, GPU and associated circuits and memory 5, 6. The device 100 can be directly connected (e.g., via WiFi or Ethernet) for conducting a teleconference with a remote client 50 via the Internet 60 or INET. The display 12 provides a user interface for operating the teleconferencing software and displaying the teleconferencing views and graphics discussed herein to the participants M1, M2...M3. Figure 1A The device or conference camera 100 may alternatively be connected directly to the Internet 60, thereby allowing the remote client 50 to record video directly to a remote server or access it in real time from such a server.
[0065] Figures 2A to 2K yes Figure 1A and Figure 1B and a schematic diagram of an embodiment of a conference camera 14 or camera tower 14 arrangement of a device or conference camera 100 suitable for capturing wide and / or panoramic scenes. Although a conference camera need not be a camera tower, "camera tower" 14 and "conference camera" 14 are used substantially interchangeably herein. Figures 2A-2K The height of the mid-width cameras 2, 3, 5 from the bottom of the device 100 is preferably greater than 8 inches and less than 15 inches.
[0066] exist Figure 2A In the camera tower 14 arrangement of FIG. 1 , a plurality of cameras are arranged circumferentially on the camera tower 14 at camera level (8 to 15 inches) and are equally spaced apart. The number of cameras is determined by the field of view of the cameras and the angle spanned, and in the case of forming a panoramic stitched view, the cumulative angle spanned should have overlap between the individual cameras. In an example Figure 2A In the case of FIG. 1 , four cameras 2a, 2b, 2c, 2d (labeled 2a-2d), each having a 100-110 degree field of view (shown in dashed lines), are arranged at 90 degrees to each other to provide a cumulative view or a stitchable view or a stitched view of 360 degrees around the camera tower 14.
[0067] In e.g. Figure 2B In the case of a tower 14, three cameras 2a, 2b, 2c (labeled 2a-2c), each having a field of view of 130 degrees or more (shown in dashed lines), are arranged 120 degrees from each other, similarly to provide a 360 degree cumulative view or stitchable view around the tower 14. The vertical field of view of cameras 2a-2d is smaller than the horizontal field of view, for example less than 80 degrees. The images, videos or sub-scenes from each camera 2a-2d can be processed to identify locations or sub-scenes of interest before or after known optical corrections such as stitching, dewarping or distortion compensation, but are typically so corrected before output.
[0068] exist Figure 2C In the camera tower 14 arrangement of FIG. 1 , a single fisheye or near-fisheye camera 3 a facing upward is arranged at the top of the camera tower 14 at camera level (8 to 15 inches). In this case, the fisheye camera lens is arranged to have a 360 degree continuous horizontal view and a vertical field of view of approximately 215 (e.g., 190-230 degrees) (shown in dashed lines). Alternatively, as Figure 2DAs shown, a single catadioptric "cylindrical image" camera or lens 3b having a cylindrical transparent shell, top parabolic mirror, black center column, telecentric lens configuration is arranged with a 360 degree continuous horizontal view, with a vertical field of view of approximately 40-80 degrees, roughly centered on the horizon. In the case of each of the fisheye and cylindrical cameras, the vertical field of view located 8-15 inches above the conference table extends below the horizon, allowing the conference participants M1, M2...Mn around the conference table to be imaged to waist level or below. The images, videos or sub-scenes from each camera 3a or 3b can be processed to identify the orientation or sub-scene of interest before or after known optical corrections for fisheye or catadioptric lenses such as dewarping or distortion compensation, but are typically so corrected before output.
[0069] exist Figure 2K In the camera tower 14 arrangement of FIG. 1 , multiple cameras are arranged circumferentially on the camera tower 14 at camera level (8 to 15 inches), equally angularly spaced. In this case, the number of cameras is not intended to form a completely continuous panoramic stitched view, and the cumulative angles spanned have no overlap between the individual cameras. In an example Figure 2K In the case of a 130 degree or higher field of view (shown in dashed lines), two cameras 2a, 2b are arranged at 90 degrees to each other to provide a separation view of about 260 degrees or higher on both sides of the camera tower 14. This arrangement will be useful in the case of a long conference table CT. In, for example, Figure 2E In the case of, the two cameras 2a-2b are panned and / or rotatable around a vertical axis to cover the locations of interest B1, B2 ... Bn discussed herein. The images, videos or sub-scenes from each camera 2a-2b may be scanned or analyzed as discussed herein before or after optical correction.
[0070] exist Figure 2F and Figure 2G In the figure, the arrangement at the head or end of the table is shown, i.e. Figure 2F and Figure 2G Each camera tower 14 shown is intended to be advantageously placed at the head of a conference table CT. Figure 3A-6A As shown, a large flat panel display FP for presentations and video conferencing is usually placed at the head or end of a conference table CT, and Figure 2F and Figure 2G The arrangement may alternatively be placed directly in front of and adjacent to the flat plate FP. Figure 2FIn the camera tower 14 arrangement, two cameras with approximately 130 degree field of view are placed 120 degrees from each other, covering both sides of the long conference table CT. The display and touch interface 12 faces under the table (especially useful if there is no tablet FP on the wall) and displays the client for the video conferencing software. This display 12 can be a connected, attachable or removable tablet or mobile device. Figure 2G In a camera tower arrangement, a high resolution, optionally tilted camera 7 (optionally connected to its own independent teleconferencing client software or instance) can be pointed towards the object of interest (e.g. a whiteboard WB or page or paper on a table CT surface), and two independent pan / tilt cameras 5a, 5b with, e.g., 100-110 degree fields of view are pointed or can be pointed to cover the locations of interest.
[0071] The images, videos or sub-scenes from each camera 2a, 2b, 5a, 5b, 7 may be scanned or analyzed as discussed herein before or after optical correction. Figure 2H It is shown that two identical units, each with two cameras 2a-2b or 2c-2d of 100-130 degrees arranged 90 degrees apart, can be used independently at the head or end of the table CT as a >180 degree view unit, but can also be optionally combined back to back to produce the same Figure 2A The same basic unit, Figure 2A The unit has four cameras 2a-2d that span the entire room and are located exactly in the middle of the conference table CT. Figure 2H Each of the tower units 14, 14 is provided with a network interface and / or a physical interface for forming a combined unit. Figure 2J , 6A , 6B and 14, the two units can be arranged freely or cooperatively alternatively or additionally.
[0072] exist Fig.2I In, similar to Figure 2C The fisheye camera or lens 3a (physically and / or conceptually interchangeable with the catadioptric lens 3b) is arranged atop the camera tower 14 at camera level (8 to 15 inches). A rotatable, high resolution, optionally tilted camera 7 (optionally connected to its own separate teleconferencing client software or instance) can be pointed at the object of interest (e.g., a whiteboard WB or page or paper on the table CT surface). Fig. 6A , Figure 6B and Fig.14 As shown, when the first telephone conference client (in Fig.14In the "conference room (local) display" or connected to the "conference room (local) display"), for example via a first physical or virtual network interface or channel 10a, the synthesized sub-scene is received from the scene SC cameras 3a, 3b as a single camera image or synthesized output CO, and the second teleconference client (in Fig.14 The arrangement works advantageously when receiving independent high resolution images from camera 7 (residing within device 100 and connected to the Internet via a second physical or virtual network interface or channel 10b).
[0073] Figure 2J A similar arrangement is shown in which similarly separate video conferencing channels for images from cameras 3a, 3b and 7 may be advantageous, but in Figure 2J In an arrangement of , each camera 3a, 3b has its own tower 14 relative to camera 7 and is optionally connected to the remaining towers 14 via an interface 15 (which may be wired or wireless). Figure 2J In an arrangement, a panoramic tower 14 with scene SC cameras 3a, 3b may be placed at the center of the conference table CT, and a directional high-resolution tower 14 may be placed at the head of the table CT or anywhere else a directional high-resolution individual client image or video stream is of interest. The images, videos or sub-scenes from each camera 3a, 7 may be scanned or analyzed as discussed herein before or after optical correction.
[0074] Use of conference cameras
[0075] refer to Figure 3A , Figure 3B and Fig.12 According to an embodiment of the present method for synthesizing and outputting a photographic scene, a device or conference camera 100 (or 200) is placed on top of, for example, a round or square conference table CT. The device 100 can be positioned according to the convenience or intention of the conference participants M1, M2, M3...Mn.
[0076] In any typical meeting, participants M1, M2...Mn will be distributed at angles relative to device 100. If device 100 is placed in the center of participants M1, M2...Mn, then the participants can be captured with a panoramic camera as discussed herein. Conversely, if device 100 is placed to one side of the participants (e.g., at one end of a table or mounted to a tablet FP), a wide camera (e.g., 90 degrees or more) may be sufficient to span participants M1, M2...Mn.
[0077] like Figure 3AAs shown, participants M1, M2, ..., Mn each have a corresponding orientation B1, B2, ..., Bn from device 100, for example measured from an origin OR for illustration purposes. Each orientation B1, B2, ..., Bn may be an angle range or a nominal angle. Figure 3B As shown, the "unfolded", projected or de-warped fisheye, panoramic or wide scene SC includes an image of each participant M1, M2 ... Mn arranged at the expected corresponding orientation B1, B2 ... Bn. In particular, in the case of a rectangular table CT and / or the device 100 is arranged at one side of the table CT, the image of each participant M1, M2 ... Mn may be perspectively foreshortened or distorted in perspective (in the case of the lateral orientation ... Figure 3B , and with the intended foreshortening direction throughout the figures). Perspective and / or visual geometry corrections known to those skilled in the art may be applied to foreshortened or perspective distorted images, sub-scenes or scenes SC, but are not required.
[0078] Face detection and widening
[0079] As an example, modern face detection libraries and APIs using common algorithms (e.g., Android's FaceDetector.Face class, Objective C's CIDetector class and CIFaceFeature objects, OpenCV's CascadeClasifier class using Haar cascades, among the more than 50 APIs and SDKs available) typically return the pupil distance, as well as the location of facial features and facial pose in space. A rough base for face width estimation can be approximately twice the pupil distance / angle, with a rough ceiling of three times the pupil distance / angle if the ears of participant Mn are included in the range. A rough base for portrait width estimation (i.e., the head plus some shoulder width) can be twice the face width / angle, with a rough ceiling of four times the face width / angle. In an alternative, a fixed angle or other more direct setting of the subscene width can be used.
[0080] Figure 4A-4B and Figure 5A-5B An exemplary two-step identification and / or separate identification of both face width and shoulder width is shown (either of which may be the minimum width used to set the initial subscene width as discussed herein). Figure 4A and 4B As shown, the face widths FW1, FW2, ... FWn are set according to the pupil distance or other size analysis of facial features (features, categories, colors, segments, blocks, textures, trained classifiers or other features) from the panoramic scene SC. Figure 5A , 5B, 6A and 6B, based on the same analysis, the shoulder widths SW1, SW2, ... SWn are set by scaling approximately 3 or 4 times, or based on the default acoustic resolution or width.
[0081] Synthesizing angle-separated sub-scenes
[0082] Fig. 7A A and B respectively show a top view and a conference camera panoramic image signal SC of a conference camera 100 use case showing a conference table CT with about ten seats, showing five participants M1, M2, M3, M4 and M5, and including a depiction of the identification of the visual minimum width Min.2 and the corresponding angular range orientation B5 of interest and the auditory minimum width Min.5 and the corresponding vector orientation B2 of interest.
[0083] exist Fig. 7A In the figure, the conference camera 100 is located in the middle of a 10-person long conference table CT. Therefore, participants M1, M2, M3 towards the middle of the table CT are least foreshortened and occupy the largest image area and angular view of the camera 100, while participants M5 and M4 towards the ends of the table CT are most foreshortened and occupy the smallest image area.
[0084] exist Figure 7B In the embodiment of the present invention, the entire scene video signal SC is, for example, a 360-degree video signal including all participants M1 ... M5. The conference table CT appears in the scene SC with a distorted "W"-shaped panoramic view feature, while the participants M1 ... M5 appear in different sizes and with different perspective foreshortening aspects (simply and schematically represented by rectangular bodies and oval heads), depending on their position and distance from the conference camera 100. Fig. 7A and 7B As shown, each participant M1...M5 can be represented in the memory 8 by a corresponding position B1...B5 determined by the auditory or visual or sensor positioning of the sound, movement or feature. Fig. 7A and 7B As shown, participant M2 can be located by detecting the face (and having a corresponding vector-like orientation B2 and a minimum width Min.2 recorded in the memory, determined in proportion to the face width derived from the face detection program), and participant M5 can be located by beamforming, relative signal strength and / or the transit time of a speech-like audio signal (and having a corresponding sector-shaped orientation B5 and a minimum width Min.5 recorded in the memory, determined in proportion to the approximate resolution of the acoustic array 4).
[0085] Fig. 8AA schematic diagram showing the panoramic video signal SC.R to be synthesized into the stage scene video signals STG, CO, the extraction of the sub-scene video signals SS2, SS5, the minimum width Min.n, and the video signal of the conference camera 100. Fig. 8A The top of the Figure 7B .like Fig. 8A As shown, the images from the image can be sorted according to the orientation (limited to orientations B2 and B5 in this example) and width (limited to widths Min.2 and Min.5 in this example) of interest. Figure 7B The overall scene video signal SC is subsampled. The sub-scene video signal SS2 is at least as wide as the (visually determined) face width limit Min.2, but may be widened or scaled wider relative to the width, height and / or available area of the stage STG or the composite output CO aspect ratio and available area. The sub-scene video signal SS5 is at least as wide as the (acoustically determined) acoustic approximation Min.5, but may be widened or scaled wider and is similarly limited. The reduced panoramic scene SC.R in this capture is a top and bottom cropped version of the overall scene SC, in this case cropped to an aspect ratio of 10:1. Alternatively, the reduced panoramic scene SC.R may be derived from the overall panoramic scene video signal SC by proportional or anamorphic scaling (e.g. the top and bottom parts are retained, but compressed more than the middle part). In any case, in Fig. 8A and 8B In the example shown, three different video signal sources SS2, SS5 and SC.R can be used for the composite stage STG or the composite output CO.
[0086] Figure 8B Basically reproduce Fig. 8A , and shows a schematic diagram of sub-scene video signals SS2, SS5 and panoramic video signal SC.R to be synthesized into a stage scene video signal STG or CO. Figures 8C to 8E Three possible composite output or stage scene video signals STG or CO are shown.
[0087] exist Figure 8CIn the composite output CO or stage scene video signal STG shown, the reduced panoramic video signal SC.R is synthesized to completely span the top of the stage STG, in this case occupying less than 1 / 5 or 20% of the stage area. Subscene SS5 is synthesized to occupy at least its minimum area, not scaled as a whole, but widened to fill approximately 1 / 2 of the stage width. Subscene SS2 is also synthesized to occupy at least its (rather small) minimum area, not scaled as a whole, and also widened to fill approximately 1 / 2 of the stage width. In this composite output CO, the two subscenes are given approximately the same area, but the participants have different apparent sizes corresponding to their distance from the camera 100. In addition, note that the left-right or clockwise order of the two subscenes as synthesized is the same as the order of the participants in the room or the orientation of interest from the camera 100 (and as it appears in the reduced panoramic view SC.R). In addition, any transfer discussed herein can be used to synthesize the subscene video signals SS2 and SS5 into the stage video signal STG. For example, two sub-scenes may simply fill the stage STG in real time; or one sub-scene may slide in from its corresponding left-right stage direction to fill the entire stage, and then gradually narrow with the help of another sub-scene sliding in from its corresponding left-right stage direction, etc. In each case, the sub-scene window, frame, outline, etc. displays its video stream through the entire transition.
[0088] exist Fig.8D In the composite output CO or stage scene video signal STG shown, the reduced panoramic video signal SC.R is similarly composited into the scene STG, but each of the signals SS5 and SS2 has been scaled or enlarged so that the participants M5, M2 occupy more of the stage STG. The minimum width of each signal SS5 and SS2 is also shown as enlarged, and the signals SS5 and SS2 still occupy no less than their corresponding minimum widths, but each signal SS5 and SS2 is widened to fill approximately 1 / 2 of the stage (in the case of SS5, the minimum width occupies 1 / 2 of the stage). The participants M5, M3 have substantially equal sizes on the stage STG or in the composite output signal CO.
[0089] exist Fig. 8E In the composite output CO or stage scene video signal STG shown, the reduced panoramic video signal SC.R is similarly composited into the scene STG, but each of the signals SS5 and SS2 is scaled or enlarged as appropriate. The sub-scene signals SS5 and SS2 still occupy no less than their respective minimum widths, but each is widened to fill a different amount of the stage. In this case, sub-scene signal SS5 is not enlarged or enlarged, but has a wider minimum width and occupies more than 2 / 3 of the stage SG. On the other hand, the minimum width of signal SS2 is depicted as being enlarged, occupying approximately 3 times its minimum width. Appearance Fig. 8E One case of the relative proportions and states of the subscene SS5 may be the following: participant M5 is not visually located, giving a wide and uncertain (low confidence level) position of interest and a wide minimum width; in addition, if participant M5 continues to speak for a long time, the share of subscene SS5 in stage STG may be optionally increased. At the same time, participant M2 may have highly reliable face width detection, allowing subscene SS2 to be scaled and / or widened to use more than its minimum width.
[0090] Fig.9A Also shown is a schematic diagram of the alternative panoramic video signal SC.R to be synthesized into the stage scene video signal, the extraction of the alternative sub-scene video signal SSn, the minimum width Min.n and the video signal of the conference camera 100. In addition to the fact that participant M1 has become the latest speaker, Fig.9A The top of the Figure 7B , where the corresponding sub-scene SS1 has the corresponding minimum width Min.1. Fig.9A As shown, the images from the image can be sorted according to the azimuths (now azimuths B1, B2, and B5) and widths (now widths Min.1, Min.2, and Min.5) of interest. Figure 7B The overall scene video signal SC is subsampled. The subscene video signals SS1, SS2 and SS5 are each at least as wide (visually, acoustically or sensor determined) as their respective minimum widths Min.1, Min.2 and Min.5, but may be widened or scaled wider relative to the width, height and / or available area of the stage STG or the composite output CO aspect ratio and available area. The reduced panoramic scene SC.R in this capture is a top, bottom and side cropped version of the overall scene SC, in this case cropped to span only the most relevant / closest speakers M1, M2 and M5, with an aspect ratio of approximately 7.5:1. Fig.9A and 9B In the example shown, four different video signal sources SS1, SS2, SS5 and SC.R can be used to composite to the stage STG or the composite output CO.
[0091] Fig. 9B Basically reproduce Fig.9A , and shows a schematic diagram of a sub-scene video signal and a panoramic video signal to be synthesized into a stage scene video signal. Figures 9C to 9E Three possible composite output or stage scene video signals are shown.
[0092] exist Fig. 9CIn the composite output CO or stage scene video signal STG shown, the reduced panoramic video signal SC.R is synthesized to almost completely span the top of the stage STG, occupying less than 1 / 4 of the stage area in this case. Subscene SS5 is synthesized again to occupy at least its minimum area, not scaled as a whole, but widened to fill about 1 / 3 of the stage width. Subscenes SS2 and SS1 are also synthesized to occupy at least their smaller minimum area, not scaled as a whole, and are also widened to fill about 1 / 3 of the stage width. In this composite output CO, the three subscenes are given roughly the same area, but the participants have different apparent sizes corresponding to their distance from the camera 100. The left-right order or clockwise order of the two subscenes such as the synthesis or transfer is the same as the order of the participants in the room or the orientation of interest from the camera 100 (and as it appears in the reduced panoramic view SC.R). In addition, any transfer discussed herein can be used to synthesize subscene video signals SS1, SS2, SS5 into the stage video signal STG. In particular, since the sliding transition is approached in or according to the same left-right order as the reduced panoramic view SC.R (e.g., if M1 and M2 are already on the stage, M5 should slide in from the right side of the stage; if M1 and M5 are already on the stage, M2 should slide in from the top or bottom between them; if M2 and M5 are already on the stage, M1 should slide in from the left side to maintain the order of M1, M2, M5 of the panoramic view SC.R), the transition is less jarring.
[0093] exist Fig.9D In the composite output CO or stage scene video signal STG shown, the reduced panoramic video signal SC.R is similarly composited into the scene STG, but each of the signals SS1, SS2 and SS5 is proportionally scaled or enlarged so that the participants M1, M2, M5 occupy more of the stage STG. The minimum width of each signal SS1, SS2, SS5 is also shown as enlarged, where the signals SS1, SS2, SS5 still occupy no less than their corresponding enlarged minimum widths, but the sub-scene SS5 is widened to fill slightly more than its enlarged minimum width on the stage, with SS5 occupying 60% of the stage width, SS2 occupying only 15%, and SS3 occupying the remaining 25%. The participants M1, M2, M5 have substantially equal heights or face sizes on the stage STG or in the composite output signal CO, although the participant M2 and the sub-scene SS2 may be significantly cropped to show only a little more than the head and / or body width.
[0094] exist Fig.9EIn the composite output CO or stage scene video signal STG shown, the reduced panoramic video signal SC.R is similarly composited into the scene STG, but each of the signals SS1, SS2, SS5 is scaled or enlarged as appropriate. The sub-scene signals SS1, SS2, SS5 still occupy no less than their respective minimum widths, but each is widened to fill a different amount of the stage. In this case, none of the sub-scene signals SS1, SS2, SS5 are enlarged or enlarged, but sub-scene SS1 with the nearest or relevant speaker M1 occupies more than 1 / 2 of the stage SG. On the other hand, each of the sub-scenes SS2 and SS5 occupies a smaller or reduced share of the stage STG, but the minimum width of sub-scene SS5 results in any further reduction of the share of the stage STG taken from sub-scene SS2 or SS1. Appearance Fig.9E One case of the relative proportions and states may be the following: the participant M1 can be visually positioned, but when the participant M1 continues to speak for a long time, the share of the stage STG of the sub-scene SS1 relative to the other two sub-scenes can be optionally increased.
[0095] exist Fig.9F In the panoramic scene SC or the reduced panoramic scene SC.R shown, the conference camera 1000 is not placed at the center of the table CT, but at one end of the table CT (for example, as shown by Fig. 7A ), where the tablet FP displays the remote conference participants. In this case, the conference table CT is again displayed as a highly distorted "W" shape. Fig.9F As shown at the top of , if the index direction or origin OR of the conference camera 100 or the panoramic scene SC is oriented so that the limitations of the high aspect ratio panoramic scene SC "segment" the conference table CT, it is quite difficult to reference the positions of the people around the table CT. However, if the index direction or origin OR of the conference camera 100 or the panoramic scene is arranged so that the table CT is continuous and / or all people are located on one side, the scene is more natural. According to this embodiment, the processor 6 can perform image analysis to change the index position or origin position of the panoramic image. In one example, the index position or origin position of the panoramic image can be "rotated" so that a single continuous segment of the image block corresponding to the table area is maximized in area (e.g., the table is not segmented). In another example, the index position or origin position of the panoramic image can be "rotated" so that the two closest or largest facial recognitions are farthest from each other (e.g., the table is not segmented). In a third example, in another example, the index position or origin position of the panoramic image can be "rotated" so that the lowest height segment of the image block corresponding to the table area is at the edge of the panorama (e.g., rotating the "W" shape to place the edge of the table closest to the conference camera 100 at the edge of the panorama).
[0096] Fig. 10A A schematic diagram showing a possible composite output CO or stage scene video signal STG, and essentially reproducing Fig.9D A composite output signal CO or stage video signal STG, in which the reduced panoramic signal is synthesized to occupy less than 1 / 4 of the top of the stage STG, and the three different sub-scene video signals are synthesized to occupy different amounts of the remaining part of the stage STG. Fig. 10B An alternative schematic diagram of a possible composite output or stage scene video signal is shown, in which three different sub-scene video signals adjacent to each other are composited to occupy different amounts of the stage STG or composite output signal CO.
[0097] Fig.11A and 11B Schematic diagram showing two alternative ways in which video conferencing software can display composite output or stage scene video signals. Fig.11A and Fig. 11B In the embodiment, the composite output signal CO is received (e.g. via a USB port) as a single camera signal with accompanying audio (optionally mixed and / or beamformed to emphasize the voice of the current speaker) and integrated into the video conferencing application as a single camera signal. Fig.11A As shown, each individual camera signal is given a separate window and a selected or active or foreground signal such as the composite output signal CO is reproduced as a thumbnail. Fig. 11B In the example shown, a selected single camera signal is given as much area on the display as is practical, and the selected or active or foreground signal, such as the composite output signal CO, is presented as a shaded thumbnail or a disabled thumbnail.
[0098] Subscene identification and synthesis
[0099] like Fig.12 As shown, in step S10, new sub-scenes SS1, SS2...SSn can be created and tracked based on the scene (e.g., based on the identification within the panoramic video signal SC). Subsequently, in step S30, the sub-scenes SS1, SS2...SSn can be synthesized based on the orientation, conditions and identification of interest discussed herein. The synthesized output or stage scene STG, CO can then be output in step S50.
[0100] exist Fig.13 In the additional details shown, as in Figures 3A to 7B (include Figure 3A and Figure 7B ), in step S12, the device 100 captures a wide-angle (eg, an angle between 90-360 degrees) scene SC with a field of view of at least 90 degrees from one or more at least partially panoramic cameras 2 or 2a...2n.
[0101] Subsequent processing for tracking and subscene identification may be performed on the native, distorted or unstitched scene SC, or may be performed on the unwrapped, distortion-corrected or stitched scene SC.
[0102] In step S14, new locations of interest Bl, B2, ... Bn are obtained from the wide-angle view SC using one or more beamforming, recognition, identification, vectoring or homing techniques.
[0103] In step S16, one or more new orientations are widened from the initial angular range (e.g., 0-5 degrees) to an angular range sufficient to span a typical human head, and / or a typical human shoulder, or other default width (e.g., measured in pixels or angular range). Note that the order of analysis can be reversed, for example, the face can be detected first, and then the orientation to the face can be determined. Widening can be performed in one, two, or more steps, with the two mentioned herein as examples; and "widening" does not need to be a gradual widening process, for example, "widening" can mean directly setting the angular range based on detection, recognition, thresholds, or values. Different methods can be used to set the angular range of the sub-scene. In some cases, such as when two or more faces are close to each other, "widening" can be selected to include all of these faces, even if only one is in the precise orientation B1 of interest.
[0104] In step S16, (and as Figure 5A and 5B ), the shoulder width subscenes SS1, SS2 ... SSn can be set or adjusted as in step S18, according to the pupil distance or a measurement obtained from other face, head, torso or other visual features (features, classes, colors, segments, blocks, textures, trained classifiers or other features), which can be obtained from the scene SC. The subscenes SS1, SS2 ... SSn width can be set according to the shoulder width (alternatively according to the face width FW), or alternatively as a predetermined width related to the angular resolution of the acoustic microphone array 4.
[0105] Alternatively, in step S16, the upper and / or lower limits of the subscene width may be set for each or all orientations of interest, or adjusted in step S18 to, for example, a peak, average, or representative shoulder width SW and face width FW, respectively. It should be noted that the symbols FW and SW are used interchangeably herein as "face width" FW or "shoulder width" SW (i.e., the span of the face or shoulder to be captured as a subscene at an angle), and the resulting face width or shoulder width subscene SS representing the face width FW or shoulder width SW (i.e., a block of pixels or subscenes with corresponding widths identified, obtained, adjusted, selected, or captured from the wide scene SC).
[0106] In step S16, or alternatively or additionally in steps S16-S18, a first discrete sub-scene (e.g., FW1 and / or SW1) of at least 20 degree angular field of view is obtained from the wide-angle scene SC at a first orientation of interest B1, B2 ... Bn. Alternatively or in addition to the at least 20 degree angular field of view (e.g., FW1 and / or SW1) setting, the first discrete sub-scene FW1 and / or SW1 can be obtained from the wide-angle scene SC as a field of view angle spanning at least 2 to 12 times the pupil distance (e.g., specific to M1 or representative of M1, M2 ... Mn), or alternatively or additionally scaled to capture a width between the pupil distance (e.g., specific to M1 or representative of M1, M2 ... Mn) and the shoulder width (e.g., specific to M1 or representative of M1, M2 ... Mn). The sub-scene capture of the wider or shoulder width SWn can record the narrower face width FWn for later reference.
[0107] If a second location of interest B1, B2 ... Bn is available, then in step S16, or alternatively or additionally in steps S16-S18, a second discrete sub-scene (e.g., FW2 and / or SS2) is obtained in a similar manner from the wide-angle view SC at the second location of interest (e.g., B2). If successive locations of interest B3 ... Bn are available, successive discrete sub-scenes (e.g., FW3 ... n and / or SS3 ... n) are obtained in a similar manner from the wide-angle view SC at successive locations of interest B3 ... Bn.
[0108] The first and second positions of interest B1, B2 (and subsequently positions of interest B3 ... Bn), whether obtained by stitching together different camera images or from a single panoramic camera, can have a substantially common angular origin to the first position of interest because they are obtained from the same device 100. Optionally, the ... can be obtained from a separate camera 5 or 7 of the device 100, or from a camera on a connected device (e.g. Figure 1A a connected laptop, tablet or mobile device 40; or Figure 2J A connected satellite camera 7 on a satellite tower 14b) obtains one or more additional locations of interest Bn starting from different angles.
[0109] As described above, the set, obtained or widened subscene SS representing the width FW or SW can be adjusted in step S18, for example (i) to have a size equal to or matching that of other subscenes; (ii) to be evenly divided or divisible (for example, into 2, 3 or 4 segments) relative to the aspect ratio of the output image or stream signal, optionally not less than the width base or more than the previously indicated maximum limit; (iii) to avoid overlapping with other subscenes near the orientation of interest; and / or (iv) to make the brightness, contrast or other video attributes match those of other subscenes.
[0110] In step S20 (which may include in a reasonable and operative combination Figure 16-18 Mode 1, 2 or 3), data and / or metadata about the identified locations of interest B1, B2...Bn and sub-scenes FW1, FW2...FWn and / or SS1, SS2...SSn may be recorded for tracking purposes. For example, the relative position, width, height and / or any adjusted parameters of the above may be recorded from an origin OR (e.g., determined by a sensor or by calculation).
[0111] Alternatively, in step S20, feature, prediction or tracking data associated with the sub-scene may be recorded, e.g. added to a sub-scene, position or other feature tracking database in step S20. For example, sub-scenes FW1, FW2 ... FWn and / or SS1, SS2 ... SSn may be instantaneous images, image blocks or video blocks identified within an image or video scene SC. In the case of video, prediction data may be associated with a scene or sub-scene, depending on the compression / decompression method of the video, and may be recorded as data or metadata associated with the sub-scene, but will tend to be part of a new sub-scene added for tracking.
[0112] After recording the trace data or other data of interest, processing returns to the main routine.
[0113] Synthesize sub-scenario for each situation
[0114] exist Fig.12In step S30, the processor 6 may synthesize the sub-scene SSn for each case (e.g., each data, flag, mark, setting or other action parameter recorded as tracking data or scene data in step S20, for example), i.e., combine the first, optionally second and optionally subsequent discrete sub-scenes SSn corresponding to different widths FW1, FW2 ... FWn and / or SW1, SW2 ... SWn into a synthetic scene or a single camera image or video signal STG or CO. In this document, a single camera image or video signal STG, CO may refer to a single frame of video or a single synthetic video frame, which represents a USB (or other peripheral bus or network) peripheral image or video signal or stream corresponding to a single USB (or other peripheral bus or network) camera.
[0115] In step S32, the device 100, its circuitry and / or its executable code may identify relevant sub-scenes SSn to be arranged in the composite, combined image or video stream STG or CO. "Relevant" may be determined according to the criteria discussed for identification in step S14 and / or updating and tracking in step S20. For example, one relevant sub-scene would be the sub-scene with the nearest speaker; and a second relevant sub-scene may be the sub-scene with the second nearest speaker. The two nearest speakers may be the most relevant until a third speaker becomes more relevant by speaking. The embodiments herein accommodate three speakers within a sub-scene within the composite scene, each with a segment of equal width or a segment of sufficient width to accommodate their head and / or shoulders. However, two speakers or four speakers or more speakers may also be easily accommodated in a wider or narrower share of the composite screen width, respectively.
[0116] By selecting subscenes SSn that enclose only the faces in height and width, up to eight speakers can be reasonably accommodated (e.g., four in the top row and four in the bottom row of the composite scene); and arrangements of from four to eight speakers can be accommodated by appropriate screen and / or window (subscene corresponding to the window) buffering and compositing (e.g., presenting the subscenes as a card with overlap, or as foreshortened view rings with more relevant speakers larger and forward and less relevant speakers smaller and backward). Fig. 6A and 6B , whenever the system determines that WB is the most relevant scene to be displayed (e.g., Fig. 6A As shown, when the auxiliary camera 7 is imaging, the scene SSn may also include whiteboard content WB. The whiteboard or whiteboard scene WB may be presented prominently, occupying most or the main part of the scene, while the speakers M1, M2...Mn or SPKR may be optionally presented together with the whiteboard WB content in a picture-in-picture manner.
[0117] In step S34, the relevant sub-scene set SS1, SS2...SSn is compared with the previously relevant sub-scene SSn. Steps S34 and S32 can be performed in reverse order. The comparison determines whether the previously relevant sub-scene SSn is available, whether it should be retained on the stage STG or CO, whether it should be removed from the stage STG or CO, whether it should be resynthesized in a smaller or larger size or perspective, or whether it needs to be changed from the previously synthesized scene or stage STG or CO in other ways. If a new sub-scene SSn should be displayed, there may be too many candidate sub-scenes SSn for the scene change. In step S36, for example, a threshold value for scene changes can be checked (this step can be performed before or between steps S32 and S34). For example, when a number of discrete sub-scenes SSn becomes greater than a threshold number (e.g., 3), the entire wide-angle scene SC or the reduced panoramic scene SC.R (e.g., is or is segmented and stacked to fit within the aspect ratio of the USB peripheral device camera) can be preferably output. Alternatively, it may be better to present a single camera scene rather than a composite scene of multiple sub-scenes SSn or as a composite output CO.
[0118] In step S38, the device 100, its circuitry and / or its executable code may set the sub-scene members SS1, SS2 ... SSn and the order in which they are transferred and / or synthesized to the synthesis output CO. In other words, having determined the candidate members of the sub-scene complements SS1, SS2 ... SSn to be output as the stage STG or CO, and whether any rules or thresholds for scene changes are met or exceeded, the order in which the scenes SSn and their transfers are added, removed, switched or rearranged may be determined in step S38. It should be noted that step S38 is more or less significantly dependent on the previous steps and the history of speakers SPKR or M1, M2 ... Mn. If two or three speakers M1, M2 ... Mn or SPKR are identified and displayed simultaneously as the device 100 begins operation, step S38 starts with a clean slate and follows the default relevant rules (e.g., present speakers SPKR in clockwise order; start with no more than three speakers in the synthesis output CO). If the same three speakers M1, M2...Mn remain relevant, then in step S38, the sub-scene membership, order and composition may not change.
[0119] As previously discussed, the identification discussed with reference to step S18 and the prediction / update discussed with reference to step S20 may result in changes to the resultant output CO in steps S32-S40. In step S40, the transfer and combination to be performed is determined.
[0120] For example, the device 100 may obtain a subsequent (e.g., third, fourth, or more) discrete sub-scene SSn at a subsequent orientation of interest from the wide-angle or panoramic scene SC. In steps S32-S38, the subsequent sub-scene SSn may be set to be synthesized or combined into a synthesized scene or a synthesized output CO. Additionally, in steps S32-S38, another sub-scene SSn other than the subsequent sub-scene (e.g., a previous or less relevant sub-scene) may be set to be removed (by synthesis transfer) from the synthesized scene (and then synthesized and output as a synthesized scene or a synthesized output CO formatted as a single camera scene) (in step S50).
[0121] As an additional or alternative example, the device 100 can set the sub-scene SSn in steps S32-S38 to be synthesized or combined into or removed from the composite scene or composite output CO according to the setting of the addition criteria or standards discussed with reference to steps S18 and / or S20 (e.g., speaking time, speaking frequency, audio cough / sneeze / doorbell, sound amplitude, voice angle, and coincidence of face recognition). In steps S32-S38, only subsequent sub-scenes SSn that meet the addition criteria can be set to be combined into the composite scene CO. In step S40, the transfer and synthesis steps to be performed are determined. Then, in step S50, the stage scene is synthesized and output as a composite output CO formatted as a single camera scene.
[0122] As an additional or alternative example, the device 100 can set the subscene SSn as a protected subscene to be protected from removal based on the retention criteria or standards (e.g., audio / speech time, audio / speech frequency, time since last speech, marked as retained) as discussed with reference to steps S18 and / or S20 in steps S32-S38. In steps S32-S38, the subscene SSn other than the subsequent subscene is removed and the protected subscene is not set to be removed from the composite scene. In step S40, the transfer and synthesis to be performed are determined. Then, in step S50, the composite scene is synthesized and output as a composite output CO formatted as a single camera scene.
[0123] As an additional or alternative example, the device 100 can set sub-scene SSn emphasis operations (e.g., zoom, flicker, genie, bounce, card sort, sort, rotation) in steps S32-S38 based on emphasis criteria or standards as discussed with reference to steps S18 and / or S20 (e.g., repeated speakers, designated presenters, recent speakers, loudest speakers, object changes rotating in the hand / scene, high-frequency scene activity in the frequency domain, raising hands). In steps S32-S38, at least one of the discrete sub-scenes SSn can be set to emphasis according to the sub-scene emphasis operations based on the corresponding or corresponding emphasis criteria or standards. In step S40, the transfer and synthesis to be performed are determined. Then, in step S50, the synthesized scene is synthesized and output as a synthesized output CO formatted as a single camera scene.
[0124] As an additional or alternative example, the device 100 may set a sub-scene participant notification or reminder operation (e.g., flashing a light at the person to the side of the sub-scene) based on a sensor or sensing criterion or standard (e.g., too quiet, remote transmission) in steps S32-S38 as discussed with reference to steps S18 and / or S20. In steps S32-S38, a local reminder flag may be set to be activated according to a notification or reminder operation based on a corresponding or corresponding sensing criterion or standard. In step S40, the transfer and synthesis to be performed are determined. Then, in step S50, the synthesized scene is synthesized and output as a synthesized output CO formatted as a single camera scene.
[0125] In step S40, the apparatus 100, its circuitry and / or its executable code generates the transitions and compositions to smoothly cause changes in the sub-scene complement of the composite image.After composition of the tracked composite output CO or other data of interest, processing returns to the main routine.
[0126] Synthetic Output
[0127] exist Fig.15In steps S52-S56 (optionally in reverse order), the synthesized scene STG or CO is formatted, i.e., synthesized, to be sent or received as a single camera scene; and / or the transition is rendered or synthesized to a buffer, screen, or frame (in this case, a "buffer", "screen", or "frame" corresponds to a single camera view output). The device 100, its circuits, and / or its executable code may use a synthesis window or screen manager, optionally with GPU acceleration, to provide an off-screen buffer for each sub-scene, and synthesize the buffer together with peripheral graphics and transition graphics into a single camera image representing a single camera view, and write the result to the output or display memory. The synthesis window or sub-screen manager circuit may perform blending, fading, scaling, rotating, copying, bending, distorting, reordering, blurring, or other processing on the buffer window, or render shadows and animations, such as flip switching, stacking switching, overlay switching, ring switching, grouping, tiling, etc. The synthesis window manager may provide visual transitions, wherein sub-scenes entering the synthesized scene may be synthesized to be added, removed, or switched with a transition effect. The sub-scenes may fade in and out, zoom in and out significantly, or radiate smoothly inward or outward. All the scenes synthesized or transferred may be video scenes, for example each comprising an uninterrupted video stream sub-sampled from the panoramic scene SC.
[0128] In step S52, the transfer or composition is rendered (repeatedly, stepwise or continuously as required) to a frame, buffer or video memory (note that the transfer and composition can be applied to individual frames or video streams, and can be an ongoing process through many video frames of the entire scene STG, CO and the individual component sub-scenes SS1, SS2...SSn).
[0129] In step S54, the device 100, its circuitry and / or its executable code may select and transfer an audio stream. Similar to the window, scene, video or sub-scene composition manager, the audio stream may be emphasized or de-emphasized, particularly in the case of the beamforming array 4, to emphasize the sub-scene being composed. Similarly, synchronizing the audio with the composed video scene may be performed.
[0130] In step S56, the device 100, its circuitry and / or its executable code outputs a simulation of the single camera video and audio as a combined output CO. As described above, this output has an aspect ratio and pixel count that simulates a single, e.g., a peripheral USB device's webcam view, e.g., an aspect ratio of less than 2:1, typically less than 1.78:1, and can be used by the group teleconferencing software as an external webcam input. When rendering the webcam input as a display view, the teleconferencing software treats the combined output CO as any other USB camera and communicates with the host device 40 (or Figure 1BAll clients that interact with a directly connected device (version 100) will be in the same directory as the host device (or Figure 1B The synthetic output CO is presented in all main views and thumbnail views corresponding to the directly connected device 100 version).
[0131] Example of subscene compositing
[0132] As reference Figure 12-16 As discussed, the conference camera 100 and the processor 6 can synthesize (in step S30) and output (in step S50) a single camera video signal STG, CO. The processor 6 operatively connected to the ROM / RAM 8 can record (in step S12) a panoramic video signal SC having an aspect ratio of substantially 2.4:1 or greater and captured from a wide camera 2, 3, 5 having a horizontal field of view of substantially 90 degrees or greater. In an optional version, the panoramic video signal has an aspect ratio of substantially 8:1 or greater, captured from a wide camera having a horizontal field of view of substantially 360 degrees.
[0133] The processor 6 may subsample (e.g., in steps S32-S40) at least two sub-scene video signals SS1, SS2, ... SSn (e.g., at the positions of interest B1, B2, ... Bn) from the wide camera 100. Figures 8C-8E and Figures 9C-9E SS2 and SS5) (e.g., in step S14). The processor 6 may synthesize (in steps S32-S40, to a buffer, frame or video memory) two or more sub-scene video signals SS1, SS2 ... SSn (e.g., in Figures 8C-8E and Figures 9C-9E SS2 and SS5) to form stage scene video signals CO, STG having an aspect ratio of substantially 2:1 or less (in steps S52-S56). Optionally, in order to densely fill the single camera video signal as much as possible (resulting in a larger view of the participant), substantially 80% or more of the area of the stage scene video signals CO, STG may be subsampled from the panoramic video signal SC. The processor 6 operably connected to the USB / LAN interface 10 may output the stage scene video signals CO, STG formatted as a single camera video signal (as in steps S52-S56).
[0134] Optimally, the processor 6 subsamples from the panoramic video signal SC (and / or optionally from a buffer, frame or video memory, such as in the GPU 6 and / or ROM / RAM 8, and / or directly from the wide cameras 2, 3, 5) additional (e.g., third, fourth or subsequent) sub-scene video signals SS1, SS2 ... SS3 (e.g., at respective locations of interest B1, B2 ... Bn) Figures 9C-9E Then, the processor may combine the two or more sub-scene video signals SS1, SS2, ... SS3 (e.g., in Figures 9C-9E SS2 and SS5) with one or more additional sub-scene video signals SS1, SS2 ... SSn (e.g., Figures 9C-9E SS1) are synthesized together to form a stage scene video signal STG, CO having an aspect ratio of substantially 2:1 or less and comprising a plurality of side-by-side sub-scene video signals (e.g., two, three, four or more sub-video signals SS1, SS2 ... SSn are synthesized in a row or a grid). It should be noted that the processor 6 can set or store in a memory one or more addition criteria for the sub-scene video signals SS1, SS2 ... SSn or one or more locations of interest. In this case, for example, only these additional sub-scene video signals SS1, SS2 ... SSn that meet the addition criteria (e.g., sufficient quality, sufficient lighting, etc.) can be transferred to the stage scene video signal STG, CO.
[0135] Alternatively or in addition, the additional sub-scene video signals SS1, SS2 ... SSn may be synthesized by the processor 6 into the stage scene video signal STG, CO by replacing one or more of the sub-scene video signals SS1, SS2 ... SSn already synthesized into the stage STG, CO to form a stage scene video signal STG, CO still having an aspect ratio of substantially 2:1 or less. Each sub-scene video signal SS1, SS2 ... SSn to be synthesized may be assigned a minimum width Min.1, Min.2 ... Min.n, and after completing each corresponding transfer to the stage scene video signal STG, CO, each sub-scene video signal SS1, SS2 ... SSn may be synthesized side by side with substantially no less than its minimum width Min.1, Min.2 ... Min.n to form the stage scene video signal STG, CO.
[0136] In some cases, for example, steps S16-S18, the processor 6 may increase the composite width of each respective sub-scene video signal SS1, SS2 ... SSn being transferred to increase during the entire transfer period until the composite width is substantially equal to or greater than the corresponding respective minimum width Min.1, Min.2 ... Min.n. Alternatively or in addition, each sub-scene video signal SS1, SS2 ... SSn may be composited side by side by the processor 6 at substantially no less than its minimum width Min.1, Min.2 ... Min.n, each SS1, SS2 ... SSn being at a respective width at which the sum of all composited sub-scene video signals SS1, SS2 ... SSn is substantially equal to the width of the stage scene video signal or composite output STG, CO.
[0137] In addition, or alternatively, the width of the sub-scene video signals SS1, SS2...SSn within the stage scene video signal STG, CO is synthesized by the processor 6 to change (for example, as in steps S16-S18) according to one or more activity criteria (for example, visual motion, sensory motion, acoustic detection of speech, etc.) detected at one or more locations of interest B1, B2...Bn corresponding to the sub-scene video signals SS1, SS2...SSn, while the width of the stage scene video signal or synthesized output STG, CO remains constant.
[0138] Optionally, the processor 6 may be configured to generate a plurality of sub-scene video signals SS1, SS2, . . . SSn by combining one or two or more sub-scene video signals SS1, SS2, . . . SSn (for example, in Figures 9C-9E SS2 and SS5) with width reduction and one or more additional or subsequent sub-scene video signals SS1, SS2 ... SSn (e.g. Figures 9C-9E SS1) by an amount corresponding to the width of one or more additional sub-scene video signals SS1, SS2 ... SSn (for example, in Figures 9C-9E SS1) is transferred to the stage scene video signal STG, CO to transfer one or more sub-scene video signals SS1, SS2...SSn (for example, in Figures 9C-9E SS2 and SS5) with one or more additional sub-scene video signals SS1, SS2 ... SSn (e.g., Figures 9C-9E SS1) are synthesized together to form a stage scene video signal.
[0139] In some cases, the processor 6 may assign a corresponding minimum width Min.1, Min.2...Min.n to each sub-scene video signal SS1, SS2...SSn, and may synthesize each sub-scene video signal SS1, SS2...SSn side by side with a width substantially not less than the corresponding minimum width Min.1, Min.2...Min.n to form a stage scene video signal or synthesized output STG, CO. When the sum of the corresponding minimum widths Min.1, Min.2...Min. of two or more sub-scene video signals SS1, SS2...SSn together with one or more additional sub-scene video signals SS1, SS2...SSn exceeds the width of the stage scene video signal STG, CO, one or more of the two sub-scene video signals SS1, SS2...SSn will be transferred by the processor 6 to be removed from the stage scene video signal or the synthesized output STG, CO.
[0140] In another alternative, the processor 9 may select at least one of two or more sub-scene video signals SS1, SS2...SSn to be transferred to be removed from the stage scene video signal STG, CO so as to correspond to a corresponding position of interest B1, B2...Bn at which one or more activity criteria (e.g., visual motion, sensory motion, acoustic detection of speech, time since the last speech, etc.) was least recently met.
[0141] In many cases, and as Figures 8B-8E and Figure 9B-9E As shown, the processor 6 can synthesize two or more sub-scene video signals SS1, SS2 ... SSn (for example, in Figures 9C-9E SS2 and SS5) with one or more additional sub-scene video signals SS1, SS2 ... SSn (e.g., Figures 9C-9E In the figure, the corresponding locations of interest B1, B2, ... Bn of SS1) are kept in the order from left to right (from top to bottom, clockwise) relative to the wide cameras 2, 3, 5.
[0142] Alternatively or in addition, the processor 6 may select each respective location of interest B1, B2 ... Bn from the panoramic video signal SC based on one or more selection criteria (e.g., visual motion, sensed motion, acoustic detection of speech, time since last speech, etc.) detected at the respective location of interest B1, B2 ... Bn relative to the wide camera 2, 3, 5. After one or more selection criteria are no longer true, the processor 6 may transfer the corresponding sub-scene video signal SS1, SS2 ... SSn for removal from the stage scene video signal or the composite output STG, CO. The selection criteria may include the presence of an activity criterion satisfied at the respective location of interest B1, B2 ... Bn. The processor 9 may calculate the time since one or more activity criteria were satisfied at the respective location of interest B1, B2 ... Bn. The processor 6 may transfer the respective sub-scene signal SS1, SS2 ... SSn for removal from the stage scene video signal STG for a predetermined period of time after one or more activity criteria were satisfied at the respective location of interest B1, B2 ... Bn.
[0143] about Figures 8A-8C , Figures 9A-9C , Fig. 10A , Figure 1B , Fig.11A , Fig. 11B and Fig. 22 The processor 6 can subsample the reduced panoramic video signal SC.R having an aspect ratio of substantially 8:1 or greater from the panoramic video signal SC. The processor 6 can then subsample the two or more sub-scene video signals (e.g., Figures 8C-8E and Figures 9C-9E SS2 and SS5) are synthesized together with the reduced panoramic video signal SC.R to form a stage scene video signal STG, CO having an aspect ratio of substantially 2:1 or less, which includes a plurality of side-by-side sub-scene video signals (e.g., Figures 8C-8E In SS2 and SS5, and in Figures 9C-9E , SS1, SS2 and SS5) and panoramic video signal SC.R.
[0144] In this case, the processor 6 may combine two or more sub-scene video signals (for example, Figures 8C-8E SS2 and SS5, and in Figures 9C-9E In the embodiment, SS1, SS2 and SS5) are synthesized together with the reduced panoramic video signal SC.R to form a stage scene video signal having an aspect ratio of substantially 2:1 or less, which includes a plurality of side-by-side sub-scene video signals (e.g., Figures 8C-8E SS2 and SS5, and in Figures 9C-9ESS1, SS2 and SS5) and a panoramic video signal SC.R above multiple side-by-side sub-scene video signals, which panoramic video signal does not exceed 1 / 5 of the area of the stage scene video signal or the composite output STG or CO, and basically extends across the width of the stage scene video signal or the composite output STG or CO.
[0145] In alternative solutions, such as Fig.24 As shown, the processor 6 may subsample or be provided with a subsample of a text video signal TD1 from a text document (e.g., from a text editor, word processor, spreadsheet, presentation, or any other document presenting text). The processor 6 may then transfer the text video signal TD1 or a rendered or simplified version TD1.R thereof into the stage scene video signal STG, CO by replacing at least one of the two or more subscene video signals with the text video signal TD1 or an equivalent TD1.R.
[0146] Optionally, the processor 6 may set one or more of the two sub-scene video signals as protected sub-scene video signals SS1, SS2 ... SSn that are protected from transfer based on one or more retention criteria (e.g., visual motion, sensed motion, acoustic detection of speech, time since last speech, etc.). In this case, the processor 6 may transfer one or more additional sub-scene video signals SS1, SS2 ... SSn to the stage scene video signal by replacing at least one of the two or more sub-scene video signals SS1, SS2 ... SSn, but specifically by transferring sub-scene video signals SS1, SS2 ... SSn other than the protected sub-scene.
[0147] Alternatively, the processor 6 may set a sub-scene emphasis operation (e.g., flashing, highlighting, outlining, icon overlay, etc.) based on one or more emphasis criteria (e.g., visual motion, sensed motion, acoustic detection of speech, time since last utterance, etc.) In this case, one or more sub-scene video signals are emphasized according to the sub-scene emphasis operation and based on the corresponding emphasis criteria.
[0148] In another variation, the processor 6 can set a sub-scene participant notification operation based on a sensing criterion from a sensor (e.g., detecting sound waves, vibrations, electromagnetic radiation, heat, UV radiation, radio, microwaves, electrical characteristics, or depth / distance detected by a sensor such as an RF element, a passive infrared element, or a range-finding element). The processor 6 can activate one or more local reminder markers based on the corresponding sensing criterion according to the notification operation.
[0149] Examples of locations of interest
[0150] For example, the location of interest may be a location corresponding to one or more audio signals or detections, such as a speaking participant M1, M2 ... Mn, which are angularly identified, vectorized or discriminated by the microphone array 4 using at least two microphones, for example, by beamforming, localization or comparable received signal strengths, or comparable transit times. Thresholds or frequency domain analysis may be used to decide whether the audio signals are strong enough or different enough, and filtering may be performed using at least three microphones to discard inconsistent pairs, multipaths and / or redundancies. The benefit of three microphones is that three pairs are formed for comparison.
[0151] As another example, alternatively or additionally, the orientation of interest can be an orientation at which motion is detected in the scene, angularly identified, vectorized or discerned by a feature, image, pattern, category and / or motion detection circuit or executable code scanning an image or motion video or RGBD from camera 2.
[0152] As another example, alternatively or additionally, the orientation of interest may be an orientation at which a facial structure is detected in the scene, angularly identified, vectorized or discerned by facial detection circuitry or executable code scanning the image or motion video or RGBD signal from camera 2. Skeletal structure may also be detected in this manner.
[0153] As another example, alternatively or additionally, the location of interest may be a location at which a substantially continuous structure of color, texture and / or pattern is detected in the scene, angularly identified, vectorized or discerned by edge detection, corner detection, peak detection or segmentation, extrema detection and / or feature detection circuitry or executable code scanning the image or motion video or RGBD signal from the camera 2. The identification may reference previously recorded, learned or trained image segments, colors, textures or patterns.
[0154] As another example, alternatively or additionally, the location of interest may be a location at which a difference from a known environment is detected in the scene, angularly identified, vectorized, or discerned by a difference and / or change detection circuit or executable code scanning the image or motion video or RGBD signal from camera 2. For example, device 100 may save one or more visual graphs of an empty conference room in which it is located, and detect when a sufficiently occluding entity (e.g., a person) occludes a known feature or area in the graph.
[0155] As another example, alternatively or additionally, the orientation of interest can be an orientation at which regular shapes such as rectangles are discerned, including a "whiteboard" shape, a door shape, or a chair back shape, as angularly identified, vectorized, or discerned by a feature, image, pattern, category, and / or motion detection circuit or executable code scanning an image or motion video or RGBD from camera 2.
[0156] As another example, alternatively or in addition, the location of interest can be a location at which a fiducial object or feature identifiable as a man-made landmark is placed by a person using device 100, including an active or passive acoustic transmitter or transducer, and / or an active or passive optical or visual fiducial marker, and / or an RFID or other electromagnetically detectable, which is angularly identified, vectorized, or distinguished by one or more of the above-described techniques.
[0157] If no initial or new position of interest is obtained in this way (e.g., because no participant M1, M2, ... Mn is still speaking), a default view can be set instead of the composite scene for output as a single camera scene. For example, as a default view, the entire panoramic scene (e.g., with a 2:1 to 10:1 H:V horizontal and vertical ratio) can be split and arranged to output a single camera ratio (e.g., typically a 1.25:1 to 2.4:1 or 2.5:1 H:V aspect ratio or horizontal and vertical ratio in the horizontal direction, although corresponding "turned" vertical ratios are also possible). As another exemplary default view before the initial position of interest is obtained, a "window" corresponding to the output scene ratio can be tracked across the scene SC at, for example, a fixed rate, for example, as a simulation of a slow panning camera. As another exemplary default view, a "headshot" of each participant M1, M2, ... Mn (plus a margin of 5-20% additional width) can be included, where the margin is adjusted to optimize the available display area.
[0158] Examples of aspect ratios
[0159] While embodiments and aspects of the invention may be useful for any angular range or aspect ratio, the benefits may optionally be greater when the sub-scenes are formed from cameras providing panoramic video signals having an aspect ratio (which aspect ratio represents frame or pixel size) of substantially 2.4:1 or greater, and are composited into a multi-participant stage video signal having an overall aspect ratio of substantially 2:1 or less (e.g., such as 16:9, 16:10, or 4:3) (as seen in most laptop or television displays (typically 1.78:1 or less), and additionally, optionally, if the stage video signal sub-scenes fill more than 80% of the composited overall frame, and / or if any further composited thumbnail forms of the stage video signal sub-scenes and the panoramic video signal fill more than 90% of the composited overall frame. In this way, each displayed participant fills nearly as much of the screen as is practicable.
[0160] The corresponding ratio between the vertical and horizontal viewing angles can be determined as a ratio according to α = 2arctan(d / 2f), where d is the vertical or horizontal dimension of the sensor and f is the effective focal length of the lens. Different wide-angle cameras used for conferencing can have a 90 degree, 120 degree, or 180 degree field of view from a single lens, and each can output a 1080p image (e.g., a 1920×1080 image) with an aspect ratio of 1.78:1 or a much wider image with an aspect ratio of 3.5:1 or other aspect ratios. When observing a conference scene, a smaller aspect ratio (e.g., 2:1 or lower) combined with a 120 degree or 180 degree wide camera can show more of the ceiling, wall, or table than might be expected. Thus, while the aspect ratio of the scene or panoramic video signal SC and the field of view FOV of the camera 100 can be independent, it may be optionally advantageous for the present embodiment to match wider cameras 100 (90 degrees or greater) with wider aspect ratio (e.g., 2.4:1 or greater) video signals, and further optionally to match the widest camera (e.g., 360 degree panoramic view) with the widest aspect ratio (e.g., 8:1 or greater).
[0161] Example of tracking subscenes or positions
[0162] Depend on Figure 1A and 1B The processing performed by the device, such as Figure 12-18 ,in particular Figure 16-18 As shown, this may include tracking sub-scenes FW, SS at locations of interest B1, B2, ... Bn within the wide video signal SC. Fig.16As shown, a processor 6 operatively connected to an acoustic sensor or microphone array 4 (with optional beam forming circuitry) and wide cameras 2, 3, 5 monitors in step S202 a substantially common angular range, which is optionally or preferably substantially 90 degrees or greater.
[0163] The processor 6 may execute code, or include or be operably connected to circuitry for identifying, in steps S204 and S206, first locations of interest B1, B2, ... Bn along a location (e.g., a measurement representing a position or orientation in Cartesian or polar coordinates, etc.) within the angular range of the wide camera 2, 3, 5, using one or both of acoustic recognition (e.g., frequency, pattern or other speech recognition) or visual recognition (e.g., motion detection, facial detection, skeleton detection, color blob segmentation or detection). As in step S10, and in steps S12 and S14, sub-sample sub-scene video signals SS from the wide camera 2, 3, 5 (e.g., newly sampled from an imaging element of the wide camera 2, 3, 5, or sub-sampled from the panoramic scene SC captured in step S12) along the locations of interest B1, B2 ... Bn identified in step S14. The width of the sub-scene video signal SS (e.g., the minimum width Min.1, Min.2...Min.n or the sub-scene display width DWid.1, DWid.2...DWid.n) can be set by the processor 6 in step S210 according to the signal characteristics of one or both of the acoustic recognition and the visual / visual recognition. The signal characteristics can represent the quality or confidence level of any of the various acoustic or visual recognitions. As used herein, "acoustic recognition" can include any recognition based on sound waves or vibrations (e.g., meeting a threshold of measurement, matching a descriptor, etc.), including frequency analysis of waveforms such as Doppler analysis, while "visual recognition" can include any recognition corresponding to electromagnetic radiation (e.g., meeting a threshold of measurement, matching a descriptor, etc.), such as heat or UV radiation, radio or microwaves, electrical characteristics recognition, or depth / range detected by sensors such as RF elements, passive infrared elements, or ranging elements.
[0164] For example, the locations of interest B1, B2, ... Bn identified in step S14 may be determined by a combination of such acoustic and visual recognition in different sequences, some of which are shown as Figure 16-18 Modes one, two or three in (which can reasonably and logically be combined with each other). In a sequence, for example, as in Fig.18 In step S220, the acoustically identified orientations are first recorded (although the order may be repeated and / or varied). Optionally, such orientations B1, B2, ... Bn may be orientations of angles, angles with tolerances, or approximate or angular ranges (e.g., Fig. 7A B5 in the figure). Fig.18As shown in steps S228-S232 of , if a sufficiently reliable visual recognition is substantially within a threshold angular range of the recorded acoustic recognition, the recorded acoustic recognition position can be refined (narrowed or re-evaluated) based on the visual recognition (e.g., facial recognition). In the same mode or in combination with another mode, for example, as in Fig.17 In step S218, any acoustic recognition not associated with visual recognition may be retained as a candidate location of interest B1, B2...Bn.
[0165] Optionally, as in Fig.16 In step S210 of , the signal characteristic represents a confidence level for either or both of the acoustic and visual recognitions. "Confidence level" need not satisfy a formal probability definition, but may represent any comparable measure that establishes a degree of reliability (e.g., crossing a threshold amplitude, signal quality, signal / noise ratio or equivalent, or success criterion). Alternatively or additionally, as in Fig.16 In step S210, the signal characteristic may represent the width of a feature identified in one or both of sound recognition (e.g., the range of angles from which the sound may originate) or visual recognition (e.g., pupil distance, face width, body width). For example, the signal characteristic may correspond to the approximate width of a face identified along the orientation of interest B1, B2...Bn (e.g., determined by visual recognition). The width of the first sub-scene video signal SS1, SS2...SSn may be set according to the signal characteristic of visual recognition.
[0166] In some cases, for example, Fig.18 In step S228, if the width is not set according to the visually recognized signal characteristics (e.g., cannot be reliably set, etc., in the case where the width defining feature cannot be identified), such as in Fig.18 In step S230, the predetermined width may be set along the acoustically recognized position detected within the angle range. Fig.18 In steps S228 and S232, if no face is recognized by image analysis along the locations of interest B1, B2, ... Bn evaluated as having acoustic signals indicative of human voice, a default width (e.g., a sub-scene having a width equal to 1 / 10 to 1 / 4 of the width of the entire scene SC) may be maintained or set, for example, as in step S230 along the acoustic locations used to define the sub-scene SS. For example, Fig. 7AA participant and speaker scene is shown, where participant M5 faces toward participant M4 and M5 is speaking. In this case, the acoustic microphone array 4 of the conference camera 100 is able to locate the speaker M5 along the orientation B5 of interest (here, the orientation B5 of interest is depicted as an orientation range rather than a vector), while image analysis of the panoramic scene SC of the wide camera 2, 3, 5 video signals may not be able to resolve the face or other visual recognition. In this case, the default width Min.5 can be set as the minimum width for initially defining, limiting or rendering the sub-scene SS5 along the orientation B5 of interest.
[0167] In another embodiment, the directions of interest B1, B2, ... Bn may be identified as acoustic recognition points detected within the angular range of the conference camera 100. In this case, the processor 6 may identify the locations close to the optional Fig.16 The visual recognition of the acoustic recognition in step S209 of the embodiment of the present invention can be consistent with the visual recognition of the acoustic recognition in step S209 (for example, within, overlapping or near the location of interest B1, B2...Bn, such as within 5-20 arcs of the location of interest B1, B2...Bn). In this case, the width of the first subscene video signal SS1, SS2...SSn can be set according to the signal characteristics of the visual recognition, which is close to or otherwise matches the acoustic recognition. This can occur, for example, in the following situation: the location of interest B1, B2...Bn is first identified using the acoustic microphone array 4, and then the location of interest B1, B2...Bn is confirmed or verified using a facial recognition that is sufficiently close or otherwise matched using the video image from the wide camera 100.
[0168] In a variation, as referenced Fig.17 and Fig.16 As described, a system including a conference camera or wide camera 100 may be used as in Fig.17 The spatial map is made using potential visual recognition or acoustic recognition as in step S218 of Fig.16 In step S209, the spatial graph is relied upon to confirm subsequent, associated, matched, close or "captured" identifications by the same or different or other identification methods. For example, in some cases, the entire panoramic scene SC may be too large to be effectively scanned on a frame-by-frame basis for facial recognition, etc. In this case, since people do not significantly move from one place to another in a meeting situation using the camera 100, especially after they are seated for the meeting, only a portion of the entire panoramic scene SC may be scanned, for example, each video frame.
[0169] For example, as in Fig.17In step S212, in order to track the sub-scenes SS1, SS2 ... SSn at the locations of interest B1, B2 ... Bn within the wide video signal, the processor 6 may scan the sub-sampling window across the motion video signal SC corresponding to the wide field of view of the camera 100 of substantially 90 degrees or greater. The processor 6 or circuitry associated therewith may identify the candidate locations of interest B1, B2 ... Bn within the sub-sampling window by substantially satisfying a threshold value defining a suitable signal quality for the candidate locations of interest B1, B2 ... Bn, for example, as Fig.17 In step S214. Each location of interest B1, B2...Bn may correspond to a location of visual recognition detected within a sub-sampling window, for example, Fig.17 In step S216. Fig.17 In step S218 of the method, the candidate locations B1, B2 ... Bn may be recorded in a spatial map (e.g., a memory or database structure that keeps track of the positions, locations and / or orientations of the candidate locations). In this way, for example, facial recognition or other visual recognition (e.g., motion) may be stored in the spatial map even if acoustic detection has not occurred at that location. Subsequently, the angular range of the wide camera 100 may be monitored by the processor 6 using an acoustic sensor or microphone array 4 for acoustic recognition (which may be used to verify the candidate locations of interest B1, B2 ... Bn).
[0170] refer to Fig. 7A For example, the processor 6 of the conference camera 100 can scan different sub-sampling windows of the entire panoramic scene SC for visual recognition (e.g., faces, colors, motion, etc.). Based on lighting, motion, orientation of faces, etc., potential locations of interest can be stored in the spatial map in FIG. 7, corresponding to the detection of faces, motion, or the like of the participants M1 ... M5. However, in Fig. 7A In the scenario shown, a potential location of interest towards participant Map.1 may not be later confirmed by acoustic signals if it corresponds to a participant who is not speaking (and this participant may never be captured in the sub-scene, but only in the panoramic scene). Once participants M1...M5 have spoken or are speaking, potential locations of interest including or towards these participants can be confirmed and recorded as locations of interest B1, B2...B5.
[0171] Optionally, as in Fig.16 In step S209, when acoustic recognition is detected close to (substantially adjacent to, near or within + / - 5-20 radians) a candidate orientation recorded in the spatial map, processor 6 may capture the orientation of interest B1, B2...Bn to correspond substantially to the one candidate orientation. Fig.16Step S209 indicates that the orientation of interest matches the corresponding part of the spatial map, and "matching" may include associating, replacing or changing the orientation of interest value. For example, because facial or motion recognition within the window and / or panoramic scene SC may have better resolution than acoustic or microphone array 4, but less frequent or less reliable detection, the detected orientations of interest B1, B2...Bn generated by sound recognition may be changed, recorded as or otherwise corrected or adjusted based on visual recognition. In this case, instead of sub-sampling the sub-scene video signals SS1, SS2...SSn along the obvious orientations of interest B1, B2...Bn derived from acoustic recognition, the processor 6 may, for example, sub-sample the sub-scene video signals along the orientations of interest B1, B2...Bn from the wide camera 100 and / or panoramic scene SC after the capture operation after correcting the acoustic orientations of interest B1, B2...Bn using previously mapped visual recognition. In this case, as in Fig.16 In step S210, the width of the sub-scene video signal SS may be set according to the detected face width or motion width, or alternatively according to the signal characteristics of the acoustic recognition (e.g., default width, resolution of the array 4, confidence level, width of features recognized in one or both of the acoustic recognition or visual recognition, approximate width of the face recognized along the orientation of interest). Fig.16 In step S210 or Fig.18 In step S230, if the sub-scene SS width is not set according to the signal characteristics of visual recognition such as face width or motion range, the predetermined width may be set according to acoustic recognition (e.g., Fig. 7A The default width in Min.5).
[0172] exist Fig.18 In the example of , the conference camera 100 and the processor 6 can track the sub-scenes at the directions of interest B1, B2...Bn by recording motion video signals corresponding to the wide camera 100 field of view FOV of substantially 90 degrees or greater. In step S220, the processor can utilize the acoustic sensor array 4 for sound recognition to monitor the angular range corresponding to the wide camera 100 field of view FOV, and when acoustic recognition is detected within the range of step S222, in step S224, the directions of interest B1, B2...Bn toward the acoustic recognition detected within the angular range can be identified. Then, the processor 6 or associated circuitry can determine the sub-scenes at the directions of interest B1, B2...Bn according to the corresponding ranges (e.g., similar to the directions of interest) B1, B2...Bn in step S226. Fig. 7AThe processor 6 may then position the sub-sampling window in the motion video signal of the panoramic scene SC, based on a range of the orientation of interest B5. If a visual recognition is detected within the range as in step S228, the processor may position the detected visual recognition within the sub-sampling window. Subsequently, the processor 6 may optionally sub-sample the sub-scene video signal SS captured from the wide camera 100 (directly from the camera 100 or the panoramic scene recorder SC) substantially centered around the visual recognition. The processor 6 may then set the width of the sub-scene video signal SS based on the signal characteristics of the visual recognition, as in step S232. In those cases where visual recognition is not possible, inappropriate, not detected, or not selected, such as in Fig.18 In step S228, the processor 6 may save or select the minimum auditory width, such as Fig.18 In step S230.
[0173] Alternatively, the conference camera 100 and the processor 6 may be connected by Figure 16-18 In, for example, Fig.17 In step S212, the acoustic sensor array 4 and the wide cameras 2, 3, 5 observing a field of view of substantially 90 degrees or more monitor the angular range to track sub-scenes at locations of interest B1, B2 ... Bn within a wide video signal such as a panoramic scene SC. The processor 6 can identify multiple locations of interest B1, B2 ... Bn, which all point to positioning within the angular range (acoustic or visual or sensor-based, as in step S216), and along with the locations of interest B1, B2 ... Bn, corresponding identifications, corresponding positioning or data representing them, such as Fig.17 In step S218, the spatial map of the recording characteristics corresponding to the positions of interest B1, B2, ... Bn is continuously stored. Fig.16 In step S210, the processor 6 can sub-sample the sub-scene video signals SS1, SS2...SSn from the wide camera 100 basically along at least one orientation of interest B1, B2...Bn, and set the width of the sub-scene video signals SS1, SS2...SSn according to the recording characteristics corresponding to at least one orientation of interest B1, B2...Bn.
[0174] Example of Predictive Tracking
[0175] In the above description of structures, devices, methods and techniques for identifying new locations of interest, various detections, identifications, triggers or other reasons for identifying such new locations of interest are described. The following description discusses updating, tracking or predicting changes in the location, direction, position, attitude, width or other characteristics of locations and subscenes of interest, and such updating, tracking and prediction may also be applied to the above description. It should be noted that the description of methods for identifying new locations of interest and updating or predicting changes in locations or subscenes is related because tracking or prediction facilitates the reacquisition of locations or subscenes of interest. The methods and techniques for identifying new locations of interest in step S14 discussed herein may be used to scan, identify, update, track, record or reacquire locations and / or subscenes in steps S20, S32, S54 or S56, and vice versa.
[0176] Predicted video data may be recorded for each sub-scene, such as data encoded according to or associated with: predictive HEVC, H.264, MPEG-4, other MPEG I slices, P slices, and B slices (or frames or macroblocks); other intra- and inter-frames, pictures, macroblocks, or slices; H.264 or other SI frames / slices, SP frames / slices (switched P), and / or multi-frame motion estimation; VP9 or VP10 super blocks, blocks, macroblocks, or super frames, intra- and inter-frame prediction, composite prediction, motion compensation, motion vector prediction, and / or partitioning.
[0177] The above-mentioned other prediction or tracking data independent of the video standard or motion compensation SPI can be recorded, for example, motion vectors derived from audio motion relative to a microphone array, or motion vectors derived from pixel-based or direct methods (e.g., block matching, phase correlation, frequency domain correlation, pixel recursion, optical flow) and / or indirect or feature-based methods (feature detection, such as corner detection with statistical functions, such as RANSAC applied to sub-scenes or scene regions).
[0178] In addition or in the alternative, the updating or tracking of each sub-scene can record, identify or score markers representing information or data or relevance thereof, such as derived audio parameters such as amplitude, speech frequency, speech length, relevant participants M1, M2...Mn (two sub-scenes with back-and-forth communication), leading or hosting participant M.Lead (a sub-scene with regular brief insertions of audio), recognized signal phrases (e.g., clapping, "keep the camera on me" and other phrases and speech recognition. These parameters or markers can be recorded independently of the tracking step or at a different time during the tracking step. The tracking of each sub-scene can also record, identify or score markers of errors or irrelevant nature, such as audio representing coughing or sneezing; regular or periodic motion or video representing machinery, wind or flickering; transient motion or motion with a frequency high enough to be transient.
[0179] Additionally or in the alternative, the updating or tracking of each sub-scene may record, identify or score a flag or data or information representing it for setting and / or protecting the sub-scene from removal, such as based on retention criteria or standards (e.g., time of audio / speech, frequency of audio / speech, time since last speech, flag for retention). In subsequent synthesis processing, removing sub-scenes other than new or subsequent sub-scenes does not remove the protected sub-scenes from the synthesized scene. That is, the protected sub-scenes have a lower priority for removal from the synthesized scene.
[0180] Additionally or alternatively, the updating or tracking of each sub-scene may record, identify or score tags or data or information representative thereof for setting additional criteria or standards (e.g., speaking time, speaking frequency, audio frequency of coughs / sneezes / doorbells, sound amplitude, voice angles, and consistency of facial recognition). In the compilation process, only subsequent sub-scenes that meet the nearby criteria can be combined into the composite scene.
[0181] Additionally or in the alternative, updating or tracking of each subscene may record, identify, or score markers for setting subscene emphasis operations based on emphasis criteria or standards (e.g., repeated speakers, designated presenters, most recent speakers, loudest speakers, motion detection of rotating objects during hand / scene changes, high frequency scene activity in the frequency domain, hand raising movements or skeleton recognition), such as audio, CGI, image, video, or synthetic effects or data or information representing the same (e.g., scaling a subscene to be larger, flashing or pulsating the boundaries of a subscene, inserting a new subscene with a sprite effect (growing from small to large), emphasizing or inserting a subscene with a bouncing effect, arranging one or more subscenes using a card sort or shuffle effect, sorting subscenes with an overlapping effect, folding back subscenes with the appearance of "folding" graphic corners). In the compilation process, at least one of the discrete subscenes is emphasized according to the subscene emphasis operation based on the corresponding or corresponding emphasis criteria.
[0182] Additionally or alternatively, the updating or tracking of each sub-scene can be based on sensors or sensed criteria (e.g., too quiet, remote transmission from social media), record, identify or score a tag or data or information representing it (e.g., flashing lights at participants M1, M2...Mn on device 100, optionally on the same side of the sub-scene) for setting sub-scene participant notification or reminder operations. In the compilation process or otherwise, based on the corresponding or corresponding sensing criteria, the local reminder tag or tags are activated according to the notification or reminder operation.
[0183] In addition or in the alternative, the updating or tracking of each sub-scene may record, identify or score a mark or data or information representing a change vector for each corresponding angular sector FW1, FW2...FWn or SW1, SW2...SWn, for example, a change in speed or direction based on each identified or located recorded characteristic (e.g., a color spot, a face, audio, as discussed herein with respect to steps S14 or S20), and / or a mark for updating the direction of the corresponding angular sector FW1, FW2...FWn or SW1, SW2...SWn based on the prediction or setting.
[0184] Additionally or alternatively, the updating or tracking of each sub-scene may record, identify or score a marker or data or information representative thereof for predicting or setting a search area for recapture or reacquisition of a lost identification or location, for example, based on the most recent position of each identified or located recorded feature (e.g., color blob, face, audio), and / or a marker for updating the direction of the corresponding angular sector based on the prediction or setting. The recorded feature may be at least one color blob, segmented or blob object representing skin and / or clothing.
[0185] Additionally or alternatively, updating or tracking of each sub-scene may maintain a Cartesian plot or specifically or optionally a polar plot of the recorded characteristic (e.g., based on an angle from an origin OR within a scene SC or an orientation B1, B2...Bn and an angular range of, for example, sub-scenes SS1, SS2...SSn corresponding to an angular sector FW / SW within the scene SC), each recorded characteristic having at least one parameter representing the orientation B1, B2...Bn of the recorded characteristic.
[0186] Thus, alternatively or additionally, embodiments of device 100, its circuitry, and / or executable code stored and executed within ROM / RAM 8 and / or CPU / GPU 6 may track sub-scenes of interest SS1, SS2 ... SSn corresponding to widths FW and / or SW within wide-angle scene SC by monitoring a target angular range (e.g., the horizontal range of cameras 2n, 3n, 5, or 7, or a subset thereof, forming scene SC) using acoustic sensor array 4 and optical sensor array 2, 3, 5, and / or 7. Device 100, its circuitry, and / or its executable code may scan target angular range SC for recognition criteria (e.g., voice, face), for example, as discussed herein with respect to step S14 (new location of interest identification) and / or step S20 (for tracking and characteristic information of locations / sub-scenes) of FIG. 8 . The device 100, its circuitry and / or its executable code may identify a first location of interest B1 based on a first identification (e.g., detection, recognition, triggering or other cause) and positioning (e.g., angle, vector, attitude or position) by at least one of the acoustic sensor array 4 and the optical sensor arrays 2, 3, 5 and / or 7. The device 100, its circuitry and / or its executable code may identify a second location of interest B2 (and optionally third and subsequent locations of interest B3 ... Bn) based on a second identification and positioning (and optionally third and subsequent identifications and positioning) by at least one of the acoustic sensor array 4 and the optical sensor arrays 2, 3, 5 and / or 7.
[0187] The device 100, its circuits and / or its executable code can set a corresponding angular sector (e.g., FW, SW, or others) for each orientation of interest B1, B2 ... Bn by expanding, widening, setting, or resetting an angular subscene (e.g., an initial small angular range or a face-based subscene FW) including the corresponding orientation of interest B1, B2 ... Bn until a threshold (e.g., reference) based on at least one identification criterion (e.g., the angle span set or reset is wider than the pupil distance, is two times or more; the angle span set or reset is wider than the head-wall contrast, distance, edge, difference, or motion transfer) is reached. Fig.13 The width threshold discussed in steps S16-S18) is met.
[0188] The device 100, its circuitry and / or its executable code may update or track (these terms are used interchangeably herein) the direction or orientation B1, B2...Bn of the corresponding angular sectors FW1, FW2...FWn and / or SW1, SW2...SWn based on changes in the direction or orientation B1, B2...Bn of the recorded characteristics (e.g., color spots, faces, audio) within or representing each identification and / or location. Optionally, as discussed herein, the device 100, its circuitry and / or its executable code may update or track each corresponding angular sector FW1, FW2...FWn and / or SW1, SW2...SWn to follow the angular changes of the first, second, and / or third and / or subsequent orientations of interest B1, B2...Bn.
[0189] Example of synthetic output (w / video conferencing)
[0190] exist Figures 8A-8D , Figures 10A-10B and Figure 19-24 In the figure, the "composite output CO", i.e. the combination of the composited and rendered / composite camera views or the composited sub-scene is shown with a lead to the main view to the remote display RD1 (representing the scene received from the conference room local display LD) and the network interface 10 or 10a, indicating that the conference room (local) display LD teleconferencing client "transparently" processes the video signal received from the USB peripheral device 100 as a single camera view and transmits the composite output CO to the remote clients or remote displays RD1 and RD2. It should be noted that all thumbnail views can also show the composite output CO. In summary, Fig.19 , 20 and 22 corresponds to Figures 3A-5B The arrangement of participants shown in Fig.21 Zhongzai Figures 3A-5B Add an additional attendee to the empty seat shown.
[0191] In an exemplary transfer, the reduced panoramic video signal SC.R (occupying approximately 25% of the vertical screen) may display a "zoomed-in" segment of the panoramic scene video signal SC (e.g., Figures 9A-9EAs shown). The zoom level can be determined by the number of pixels contained in approximately 25%. When a person / object M1, M2...Mn becomes relevant, the corresponding sub-scene SS1, SS2...SSn is transferred (for example, by a synthetic sliding video panel) to the stage scene STG or the synthetic output CO, maintaining its clockwise or left-to-right position in the participant M1, M2...Mn. At the same time, a processor using the GPU 6 memory or ROM / RAM 8 can slowly scroll the reduced panoramic video signal SC.R to the left or right to display the current location of interest B1, B2...Bn in the center of the screen. The current location of interest can be highlighted. When a new relevant sub-scene SS1, SS2...SSn is identified, the reduced panoramic video signal SC.R can be rotated or translated so that the nearest sub-scene SS1, SS2...SSn is highlighted and located in the center of the reduced panoramic video signal SC.R. With this configuration, during the conference, the reduced panoramic video signal SC.R is continuously re-rendered and virtually translated to display the relevant part of the room.
[0192] like Fig.19 As shown, in a typical video conference display, each participant's display shows a main view and multiple thumbnail views, all of which are basically determined by the output signal of the network camera. The main view is usually one of the remote participants, and the thumbnail views represent other participants. Depending on the video conference or chat system, the main view can be selected to show the active speaker among the participants, or it can be switched to another participant, usually by selecting a thumbnail, including the local scene in some cases. In some systems, the local scene thumbnail is always retained in the entire display range, so that each participant can position himself relative to the camera to present a useful scene (this example is shown in Fig.19 ).
[0193] like Fig.19 As shown, embodiments of the present invention provide a composite stage view of multiple participants rather than a single camera view. Fig.19 , potentially interesting locations B1, B2, and B3 for participants M1, M2, and M3 (represented by icons M1, M2, and M3) are available to conference camera 100. As described herein, because there are three possible participants M1, M2, and M3 located or otherwise identified and one SPKR is speaking, stage STG (equivalent to composite output CO) may initially be occupied by a default number (two in this case) of relevant sub-scenes, including the active speaker SPKR ( Fig.19 Participant M2).
[0194] Fig.191 shows three participant displays: a local display LD, such as a personal computer attached to a conference camera 100 and the Internet INET; a first personal computer ("PC") or flat panel display remote display RD1 of a first remote participant A.hex, and a second PC or flat panel display RD2 of a second remote participant A.diamond. As expected in a video conferencing environment, the local display LD primarily displays the remote speaker ( Fig.19 A.hex in ), while the two remote displays RD1, RD2 display the view selected by the remote operator or software (e.g., the view of the active speaker, the composite view CO of the conference camera 100).
[0195] Although the arrangement of participants in the main view and thumbnail view depends to some extent on user selection or even automatic selection within the video conferencing or video chat system, Fig.19 In the example of , the local display LD typically displays a main view in which the last selected remote participant is shown (e.g., A.hex, the participant working with the PC or laptop with the remote display RD1) and a row of thumbnails in which substantially all participants are represented (including the composite stage view from the local conference camera 100). In contrast, the remote displays RD1 and RD2 both display a main view including the composite stage view CO, STG (e.g., because the speaker SPKR is currently speaking), wherein this row of thumbnails also contains the remaining participant views.
[0196] Fig.19 Assume that participant M3 has spoken, or has been pre-selected as the default occupant of stage STG, and has occupied the most relevant sub-scene (e.g., the most recently relevant sub-scene). Fig.19As shown, the sub-scene SS1 corresponding to the speaker M2 (icon M2, in the remote display 2, the outline of the open mouth M2) is synthesized into a single camera view with a sliding transition (indicated by the box line arrow). The preferred sliding transition starts with zero or negligible width, in the middle, that is, the interesting positions B1, B2...Bn of the corresponding sub-scenes SS1, SS2...SSn slide onto the stage, and then the width of the synthesized corresponding sub-scenes SS1, SS2...SSn increases until at least the minimum width is reached, and the width of the synthesized corresponding sub-scenes SS1, SS2...SSn can continue to increase until the entire stage is filled. Because the synthesized (intermediate transition) and synthesized scenes are provided as camera views to the conference room (local) display LD of the teleconference client, the synthesized and synthesized scenes can be presented in the main view and thumbnail of the local client display LD and the two remote client displays RD1, RD2 at essentially the same time (i.e., presented as the current view).
[0197] exist Fig. 20 in Fig.19 Afterwards, participant M1 becomes the closest and / or most relevant speaker (e.g., the previous case is Fig.19 , where participant M2 is the nearest and / or most relevant speaker). Subscenes SS3 and SS2 of participants M3 and M2 remain relevant according to the tracking and recognition criteria and can be re-synthesized to a smaller width as needed (by scaling or cropping, optionally limited to a width limit of 2-12 times the pupil distance, as discussed herein). Subscene SS2 is similarly synthesized to a compatible size and then synthesized onto stage STG via a slide transfer (again indicated by the box line arrow). As described herein with respect to FIG. 9, Figures 10A-10B , Figures 11A-11B As described, because the new speaker SPKR is participant M1 who is azimuthally to the right (from a top to bottom perspective, clockwise) of the already displayed participant M2, sub-scene SS1 can optionally be transferred to the stage in a manner that maintains the left-to-right handedness or sequence (M3, M2, M1), in this case from the right.
[0198] exist Fig.21 in Fig. 20Thereafter, the new participant M4 who arrives in the room becomes the closest and most relevant speaker. Subscenes SS2 and SS1 of speakers M2 and M1 remain relevant according to the tracking and identification criteria and remain composited as a "3 to 1" width. The subscene corresponding to speaker M3 is "eliminated" and is no longer as relevant as the closest speaker (although many other priorities and relevances are described herein). Subscene SS4 corresponding to speaker M4 is composited to a compatible size and then composited to the camera output via a flip transfer (again represented by the box line arrow), with subscene SS3 flipped and removed. This can also be a sliding or alternative transfer. Although not shown, as an alternative, since the new speaker SPKR is participant M4 to the left (from a top to bottom perspective, clockwise) of the already displayed participants M2 and M1, subscene SS4 can be optionally transferred to the stage in a manner that maintains the left-to-right handedness or order (M4, M2, M1), in this case by transferring from the left side. In this case, sub-scenes SS2 and SS1 can each be shifted one position to the right, and sub-scene M3 can exit (slide away) the right side of the stage.
[0199] As stated in this article, Figure 19-21 An exemplary local and remote video conferencing mode, for example on a mobile device, is shown, where a composite scene that is synthesized, tracked and / or displayed has been received and displayed as a single camera scene. These are mentioned and described in the context of the previous paragraphs.
[0200] While the overall message is similar, Fig. 22 A form of video conferencing is presented, which is Fig.19 In particular, although Fig.19 In , thumbnails do not overlap the main view, and thumbnail views that match the main view are kept in the thumbnail row, but in Fig. 22 In the form of, the thumbnails overlap with the main view (e.g., are synthesized to be superimposed on the main view), and the current main view is de-emphasized in the thumbnail row (e.g., by being darkened, etc.).
[0201] Fig.23 Shows Figure 19-22 A variant in which a fourth client corresponding to a high resolution, close-up or just a single camera 7 has its own client connected to the teleconference group via network interface 10b, while the composite output CO and its transfer are presented to the conference room (local) display LD via network interface 10a.
[0202] Fig.24 Show Figure 19-22A variation of , in which a code or document review client with a text review window is connected to the conference camera 100 via a local wireless connection (although in one variation, the code or document review client can be connected from a remote station via the Internet). In one example, a first device or client (PC or tablet) runs a video conferencing or chat client that displays the attendees in a panoramic view, and a second client or device (PC or tablet) runs a code or document review client and provides it to the conference camera 100 as a video signal in the same form as the webcam. The conference camera 100 synthesizes the document window / video signal of the code or document review client onto the stage STG or CO as a full-frame sub-scene SSn, and can also optionally synthesize a local panoramic scene including the conference attendees, such as the stage STG or CO above. In this way, the text shown within the video signal can be used for all participants instead of the individual attendee sub-scenes, but the attendees can still be noted by referring to the panoramic view SC. Although not shown, the conference camera 100 device can alternatively create, instantiate or execute a second video conferencing client to host the document view. Alternatively, a high resolution, close-up or just a single camera 7 has its own client connected to the teleconference group via network interface 10b, while the composite output CO and its transfer is presented to the meeting room (local) display via network interface 10a.
[0203] In at least one embodiment, the participants M1, M2, ... Mn may be always displayed in the stage scene video signal or the composite output STG, CO. Fig.25 As shown, for example, based on at least the face width detection, the processor 6 can crop the face as a face-only sub-scene SS1, SS2...SSn and arrange it along the top or bottom of the stage scene video signal or the composite output STG, CO. In this case, it is expected that the participant using a device such as a remote device RD1 can click or touch (in the case of a touch screen) the cropped face-only sub-scene SS1, SS2, SSn to communicate with the local display LD to create a stage scene video signal STG centered on the person. In an exemplary solution, using something like Figure 1B And with the configuration of connecting directly to the Internet INET, the conference camera 100 can create or instantiate an appropriate number of virtual video conference clients and / or assign a virtual camera to each.
[0204] Fig.26Some of the diagrams and symbols used throughout Figures 1-26 are shown. In particular, arrows extending from the center of the camera lens can correspond to locations of interest B1, B2...Bn, regardless of whether the arrows are so labeled in each view. Dashed lines extending from the camera lens at an open "V" angle can correspond to the field of view of the lens, regardless of whether the dashed lines are so labeled in each view. A brief "stick figure" depiction of a person with an oval head with a square or trapezoidal body can correspond to a conference participant, regardless of whether the simplified person is so labeled in each view. A depiction of an open mouth on a simplified person can depict the current speaker SPKR, regardless of whether the simplified person with an open mouth is so labeled in each view. A wide arrow extending from left to right, right to left, top to bottom, or in a spiral shape can indicate an ongoing transition or a synthesis of transitions, regardless of whether the arrow is so labeled in each view.
[0205] In this disclosure, "wide angle camera" and "wide scene" depend on the field of view and the distance to the subject, and include any camera with a field of view wide enough to capture two different people in a meeting who are not shoulder to shoulder.
[0206] Unless a vertical field of view is specified, the "field of view" is the horizontal field of view of the camera. As used herein, a "scene" means an image (still or moving) of a scene captured by a camera. Typically, although not without exception, a panoramic "scene" SC is one of the largest images or video streams or signals processed by the system, whether the signal is captured by a single camera or stitched from multiple cameras. The scene "SC" most often referred to herein includes a scene SC, which is a panoramic scene SC captured by an equiangular distribution of cameras coupled to a fisheye lens, a camera coupled to a panoramic optical device, or overlapping cameras. The panoramic optical device can basically provide a panoramic scene directly to the camera; in the case of a fisheye lens, the panoramic scene SC can be a horizontal strip, where the field of view or horizontal strip of the fisheye view has been isolated and dewarped to a long high aspect ratio rectangular image; and in the case of overlapping cameras, the panoramic scene can be stitched and cropped (and possibly dewarped) from individual overlapping views. "Subscene" refers to a sub-portion of a scene, such as a continuous and generally rectangular block of pixels that is smaller than the entire scene. A panoramic scene can be cropped to less than 360 degrees and still referred to as an entire scene SC in which a subscene is processed.
[0207] As used herein, "aspect ratio" is discussed as the H:V horizontal:vertical ratio, where a "larger" aspect ratio increases the ratio of the horizontal relative to the vertical (wide and short). Aspect ratios greater than 1:1 (e.g., 1.1:1, 2:1, 10:1) are considered "landscape format", and for purposes of this disclosure, aspect ratios equal to or less than 1:1 are considered "portrait format" (e.g., 1:1.1, 1:2, 1:3). A "single camera" video signal is formatted as a video signal corresponding to one camera, such as UVC, also referred to as "USB Device Class Definition for Video Devices" 1.1 or 1.5 by the USB Developers Forum, each of which is incorporated herein by reference in its entirety (see, http: / / www.usb.org / developers / docs / devclass_docs / USB_Video_Class_1_5.zip USB_Video_Class_1_1_090711.zip at the same URL). Any signal discussed in UVC may be a "single camera video signal", regardless of whether that signal is transmitted, carried, transferred, or tunneled over USB.
[0208] "Display" means any direct display screen or projection display. "Camera" means a digital imager, which can be a CCD or CMOS camera, a thermal imaging camera, or an RGBD depth or time-of-flight camera. The camera can be a virtual camera formed by two or more stitched camera views and / or have a wide aspect ratio, panoramic, wide angle, fisheye, or catadioptric perspective.
[0209] A "participant" is a person, device, or location connected to a group video conferencing session and showing a view from a webcam; although in most cases a "conference attendee" is a participant who is also in the same room as conference camera 100. A "speaker" is a participant who is currently speaking or has spoken recently enough for conference camera 100 or an associated remote server to recognize him or her; but in some descriptions it may also be a participant who is currently speaking or has spoken recently enough for a video conferencing client or an associated remote server to recognize him or her.
[0210] "Compositing" generally means digital compositing as known in the art, that is, digitally assembling multiple video signals (and / or images or other media objects) to produce a final video signal, including techniques such as alpha compositing and blending, anti-aliasing, node-based compositing, keyframes, layer-based compositing, nested compositing or typography, deep image compositing (using color, opacity, and depth using depth data, whether function-based or sample-based). Compositing is an uninterrupted process, including the movement and / or animation of sub-scenes that each contain a video stream, for example, different frames, windows, and sub-screens in the entire stage scene can display different uninterrupted video streams as they are moved, transferred, blended, or otherwise composited into the entire stage scene. Compositing as used herein can use a compositing window manager or a stacking window manager with one or more off-screen buffers for one or more windows. Any off-screen buffer or display memory content can be double or triple buffered or otherwise buffered. Composition may also include processing on either or both of the buffer or display memory windows, such as applying 2D and 3D animation effects, blending, fading, scaling, enlarging, rotating, copying, bending, twisting, rearranging, blurring, adding shadows, glowing, previewing, and animation. It may include applying these to vector-oriented graphic elements or pixel-oriented or voxel-oriented graphic elements. Composition may include rendering pop-up previews after touch, mouse hover, hover, or click, window switching by rearranging several windows relative to the background to allow selection by touch, mouse hover, hover, or click, and flip switching, overlay switching, ring switching, Exposé switching, etc. As discussed herein, various visual transitions may be used on the stage - fade out, slide, grow or shrink, and combinations of these. "Transfer" as used herein includes the necessary compositing steps.
[0211] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be directly embodied in hardware, embodied in a software module executed by a processor, or a combination of the two. The software module may reside in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium may be coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. In an alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and the storage medium may reside in a user terminal as discrete components.
[0212] All of the above processes may be embodied in and fully automated by software code modules executed by one or more general or special computers or processors. The code modules may be stored on any type of computer-readable medium or other computer storage device or collection of storage devices. Some or all of the methods may alternatively be embodied in special-purpose computer hardware.
[0213] All methods and tasks described herein can be performed and fully automated by a computer system. In some cases, a computer system may include multiple different computers or computing devices (e.g., physical servers, workstations, storage arrays, etc.) that communicate and interact to perform the described functions through a network. Each such computing device typically includes a processor (or multiple processors or circuits or circuit sets, such as modules) that executes program instructions or modules stored in a memory or other non-temporary computer-readable storage medium. Various functions disclosed herein can be embodied in such program instructions, although some or all of the disclosed functions can be implemented in a dedicated circuit (e.g., ASIC or FPGA) of a computer system alternately. In the case where a computer system includes multiple computing devices, these devices can be but not necessarily located in the same location. The results of the disclosed methods and tasks can be continuously stored by transferring physical storage devices such as solid-state memory chips and / or disks into different states.
Claims
1. A method for tracking a sub-scene at a location of interest within a wide video signal, comprising: Monitoring angular ranges using acoustic sensor arrays and wide cameras observing a field of view of approximately 90 degrees or more; identifying a first location of interest along a first position within the angular range; identifying a second location of interest along the second location; subsampling a first subscene video signal from the wide camera along the first azimuth of interest; subsampling a second subscene video signal from the wide camera along the second azimuth of interest; In response to determining that a first confidence level for the first location is lower than a second confidence level for the second location, setting a first width of the first subscene video signal to be wider than a second width of the second subscene video signal; as well as The first sub-scene video signal and the second sub-scene video signal are synthesized into a stage scene video signal, wherein the first sub-scene video signal fills a larger stage scene width of the stage scene video signal than the second sub-scene video signal.
2. The method according to claim 1, wherein: Each of the first sub-scene video signal and the second sub-scene video signal is at least as wide as a minimum width limit set according to the first confidence level and the second confidence level.
3. The method according to claim 1, wherein: A first width of the first sub-scene video signal synthesized into the stage scene video signal is set according to a first confidence level from acoustic sensing, and a second width of the second sub-scene video signal is set according to a second confidence level from visual sensing, the first sub-scene video signal being wider than the second sub-scene video signal.
4. The method according to claim 1, wherein: A first width of the first sub-scene video signal synthesized into the stage scene video signal is set based on a first confidence level in the absence of visual sensing, and a second width of the second sub-scene video signal is set based on a confidence level from visual sensing, the first sub-scene video signal being wider than the second sub-scene video signal.
5. The method according to claim 1, wherein: The second sub-scene video signal synthesized into the stage scene video signal is gradually widened over time depending on acoustic detection, thereby increasing the share of the second sub-scene video signal in the stage scene video signal.
6. The method according to claim 1, wherein: Combining the first subscene video signal and the second subscene video signal includes combining the first subscene video signal and the second subscene video signal in a left-right order or a clockwise order based on the first interesting orientation and the second interesting orientation.
7. A device for tracking a sub-scene at a location of interest within a wide video signal, the device comprising: Acoustic sensor arrays; wide camera; as well as At least one processor configured to: monitoring an angular range using said acoustic sensor array and said wide camera observing a field of view of substantially 90 degrees or greater; identifying a first location of interest along a first position within the angular range; identifying a second location of interest along the second location; subsampling a first subscene video signal from the wide camera along the first azimuth of interest; subsampling a second subscene video signal from the wide camera along the second azimuth of interest; In response to determining that a first confidence level for the first location is lower than a second confidence level for the second location, setting a first width of the first subscene video signal to be wider than a second width of the second subscene video signal; as well as The first sub-scene video signal and the second sub-scene video signal are synthesized into a stage scene video signal, wherein the first sub-scene video signal fills a larger stage scene width of the stage scene video signal than the second sub-scene video signal.
8. The device according to claim 7, wherein: Each of the first sub-scene video signal and the second sub-scene video signal is at least as wide as a minimum width limit set according to the first confidence level and the second confidence level.
9. The device according to claim 7, wherein: A first width of the first sub-scene video signal synthesized into the stage scene video signal is set according to a first confidence level from acoustic sensing, and a second width of the second sub-scene video signal is set according to a second confidence level from visual sensing, the first sub-scene video signal being wider than the second sub-scene video signal.
10. The device according to claim 7, wherein: A first width of the first sub-scene video signal synthesized into the stage scene video signal is set based on a first confidence level in the absence of visual sensing, and a second width of the second sub-scene video signal is set based on a confidence level from visual sensing, the first sub-scene video signal being wider than the second sub-scene video signal.
11. The device according to claim 7, wherein: The second sub-scene video signal synthesized into the stage scene video signal is gradually widened over time depending on acoustic detection, thereby increasing the share of the second sub-scene video signal in the stage scene video signal.
12. The device according to claim 7, wherein: Combining the first subscene video signal and the second subscene video signal includes combining the first subscene video signal and the second subscene video signal in a left-right order or a clockwise order based on the first interesting orientation and the second interesting orientation.
Citation Information
Patent Citations
System and method for distributed meetings
US20040263636A1
Video conferencing
US20080218582A1
Scalable video encoding in a multi-view camera system
US20100157016A1
Tiering and manipulation of peer's heads in a telepresence system
WO2014168616A1