Video synthesis system, video synthesis method, and program

The video compositing system enhances facial expression visibility in multi-participant video conferencing by using multiple imaging devices to generate and synthesize clear close-up videos, addressing the challenges of camera placement and angle limitations.

JP2025117851APending Publication Date: 2025-08-13RICOH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024012804
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2025-08-13

AI Technical Summary

Technical Problem

Existing video conferencing systems struggle to clearly capture and display the facial expressions of multiple participants due to limitations in camera placement and angle, making it difficult to generate a composite image with clear facial expressions for all participants.

Method used

A video compositing system utilizing multiple imaging devices, each generating and synthesizing close-up videos of participants to eliminate overlaps and ensure clear facial expressions, using a combination of camera modules, image processing units, and synthesis units to generate a composite close-up video.

Benefits of technology

Improves the visibility of facial expressions of multiple participants by ensuring all participants' faces are clearly visible in the composite video, even in large conference rooms where participants may be at different distances and angles from the camera.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025117851000001_ABST
    Figure 2025117851000001_ABST
Patent Text Reader

Abstract

To improve visibility of expression of each of a plurality of persons included in video.SOLUTION: There is provided a video synthesis system including a plurality of imaging apparatuses. The imaging apparatus includes: a generation unit for generating a first close-up video including partial videos of a predetermined number of persons extracted from videos imaged by the imaging apparatus, the partial video being respectively scaled into a predetermined size; and a synthesis unit for generating third close-up video including videos of the predetermined number of persons, on the basis of the first close-up video and second close-up video including respective partial videos of a predetermined number of persons extracted from the videos imaged by another imaging apparatus, the respective partial video being scaled in a predetermined size. When the partial video of the same person is included in the first close-up video and the second close-up video, the synthesis unit makes the partial video of any one of the first close-up video and the second close-up video be included in the third close-up video.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image synthesis system, an image synthesis method, and a program. [Background technology]

[0002] In imaging devices such as cameras used in video conferencing, there is known a technology for expanding an ultra-wide-angle (fisheye) image captured using an ultra-wide-angle (fisheye) lens into a rectangular panoramic image (360° panoramic image in the case of a fisheye).

[0003] In the case of video conferencing, panoramic video allows all participants to be photographed and monitored simultaneously using a single imaging device.

[0004] Furthermore, a technique for displaying a close-up of a speaker from within a panoramic image using voice beamforming or person (face) detection means is already known.

[0005] By zooming in on the speaker, participants from other locations can also see who is speaking, and the close-up enlarged display makes it easier to see the speaker's facial expression.

[0006] Furthermore, a new technology has recently been developed that improves the sense of realism of video conferences by providing a multi-stream video layout with close-ups of all participants' faces, rather than broadcasting a video of the entire conference venue. This allows the video to be broadcast to other locations, showing the facial expressions of all participants regardless of whether they are speaking. For example, in the video shown in Figure 1, area a1 is a panoramic image that includes all participants at one location in a video conference between two locations. Areas a2 to a5 are close-up images of each participant included in area a1. Area a2 is an image of a participant at the other location.

[0007] However, in order to generate footage that shows close-ups of all meeting participants and clearly shows their facial expressions, certain conditions must be met, such as the meeting space being large enough for all meeting participants to be seated near one imaging device, and all participants facing the direction of one imaging device.

[0008] Therefore, in a large conference room, it is difficult to capture the facial expressions of everyone using a single imaging device due to issues such as the distance from the imaging device to each participant and the angle between the imaging device and each participant, which reduces the flexibility in where the imaging device is installed.

[0009] Patent document 1 describes that by placing a video conference terminal near each of multiple users who gather at one location to hold a video conference, and by combining images taken by each video conference terminal, clear images of each user can be output at other locations. Summary of the Invention [Problem to be solved by the invention]

[0010] However, with the technology of Patent Document 1, when multiple users are shown on one video conference terminal (image capture device), it is difficult to create a final composite image in which the facial expressions of all participants are clearly visible.

[0011] The present invention has been made in view of the above points, and has as its object to improve the visibility of facial expressions of each person in a video including multiple people. [Means for solving the problem]

[0012] In order to solve the above problem, the present invention provides a video compositing system including a plurality of imaging devices, each of which has a generation unit that generates a first close-up video including partial images of a predetermined number of people extracted from an image captured by the imaging device and scaled to a predetermined size, and a synthesis unit that generates a third close-up video including images of the predetermined number of people based on the first close-up video and a second close-up video including partial images of the predetermined number of people extracted from an image captured by another of the imaging devices and scaled to a predetermined size, and when the first close-up video and the second close-up video include partial images of the same person, the synthesis unit includes either of the partial images in the third close-up video. [Effects of the Invention]

[0013] This makes it possible to improve the visibility of the facial expressions of each person in a video that includes multiple people. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 10 is a diagram showing an example of a multi-stream video layout. [Figure 2] FIG. 10 is a diagram illustrating an example of the configuration of a linking function of a video conference terminal realized by utilizing a video synthesis function in an embodiment of the present invention. [Figure 3] 2 is a diagram illustrating an example of a functional configuration of a conference terminal 100 according to an embodiment of the present invention. FIG. [Figure 4] FIG. 10 is a diagram illustrating a close-up image. [Figure 5] FIG. 2 is a diagram illustrating an example of the configuration of a power supply management unit 60. [Figure 6] FIG. 10 is a diagram for explaining a specific situation in the present embodiment. [Figure 7] 10 is a diagram showing an example of a close-up image before composition generated in each conference terminal 100. FIG. [Figure 8] FIG. 10 is a diagram for explaining an example of a synthesis process in the conference terminal 100-2. [Figure 9] FIG. 10 is a diagram for explaining an example of a synthesis process in the conference terminal 100-1. [Figure 10] 10 is a flowchart illustrating an example of a processing procedure executed by the conference terminal 100. [Figure 11] 10 is a flowchart illustrating a window selection process. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. FIG. 2 is a diagram showing an example of the configuration of a linking function of a video conferencing terminal, which is realized by utilizing a video synthesis function in an embodiment of the present invention. A video conference is a type of communication such as a conference in which two or more locations exchange video including audio, allowing each location to understand the situation at the other locations through video. A video conference may be a television conference or a web conference.

[0016] 2 shows a system configuration in which a video conference is held between two locations, location a and location b. A location refers to a space, such as a conference room, where participants in the video conference (hereinafter simply referred to as "participants") gather. Location a includes a host device 200 and multiple conference terminals 100, such as conference terminals 100-1 to 100-N. The multiple conference terminals 100 at each location are an example of a video synthesis system. Furthermore, each conference terminal 100 is an example of an imaging device.

[0017] The host device 200 is a device that has a video conferencing app such as Teams (registered trademark) or Zoom (registered trademark) installed and has communication and UI functions. The host device 200A transmits videos captured by the conference terminals 100 at site a and audio collected by the conference terminals 100 to the host device 200B at site b, and controls the display of videos and output of audio at site b from the host device 200B. For example, a PC (Personal Computer) or an IWB (Interactive White Board: an electronic whiteboard with a blackboard function that allows mutual communication) may be used as the host device 200.

[0018] One host device 200 may be installed at each location. In FIG. 2, the identifier of each location ("A" or "B") is added to the end of the reference number of the host device 200 at that location. Note that FIG. 2 shows an example in which the host devices 200 at two locations are connected to the server 300 via the network 400, and communications between the host devices 200 at the two locations are mediated via the server 300. However, the host devices 200 at three or more locations may also be connected to the server 300 via the network 400. In this case, communications between the host devices 200 at any of the three or more locations are mediated by the server 300, thereby enabling a video conference to be held between any of the multiple locations.

[0019] The conference terminal 100 is a terminal (device) that has hardware and software necessary for video conferencing, such as a camera, a microphone, and a speaker. The conference terminal 100 transmits video captured by the camera and audio collected by the microphone to the host device 200. The conference terminal 100 also outputs audio (from other locations) transmitted from the host device 200 from a speaker 422.

[0020] Each conference terminal 100 functions as a USB device (has USB device functionality) and also has USB host functionality. Therefore, each conference terminal 100 can connect to other conference terminals 100 using the USB host functionality. That is, multiple conference terminals 100 can be electrically connected in series. By connecting one end of multiple serially connected conference terminals 100 to the host device 200, the multiple conference terminals 100 can be daisy-chained (connected in series) to the host device 200. In this embodiment, the conference terminal 100 closest to the host device 200 (directly connected to the host device 200) is considered to be the most significant, and the conference terminal 100 farthest from the host device 200 is considered to be the least significant. The conference terminal 100 connected to the most significant side of a certain conference terminal 100 is considered to be the most significant conference terminal 100, and the conference terminal 100 connected to the least significant side is considered to be the least significant. Video is transferred from the least significant (one end) conference terminal 100 to the most significant (other end) conference terminal 100. Furthermore, power is supplied to each of the conference terminals 100 from a single power supply system.

[0021] Each of the multiple conference terminals 100 is arranged within a site so that all of the conference terminals 100 can capture images of all participants. Therefore, the number of conference terminals 100 arranged may be changed depending on the size of the site. Each conference terminal 100 captures images of multiple participants, which are a portion of all participants, and generates a close-up image of a predetermined number of participants from the multiple participants (hereinafter referred to as a "close-up image"). Each conference terminal 100 combines the close-up image generated by a lower-level conference terminal 100 with the close-up image generated by itself, and transmits the combined close-up image to the higher-level conference terminal 100. At this time, each conference terminal 100 does not simply combine the close-up image received from the lower-level conference terminal 100 with the close-up image generated by itself, but rather combines the images so that overlaps of the same participant are eliminated. If there are multiple conference terminals 100 (imaging devices), it is possible to capture an image of one participant from different angles, making it easier to select an image that is closer to the front of each participant.

[0022] Therefore, the host device 200 receives a close-up video of participants without overlapping from the top-level conference terminal 100. The host device 200 transmits the video to host devices 200 at other locations.

[0023] By appropriately arranging multiple conference terminals 100, it is possible to display close-ups of participants far from the host device 200 in a large conference room, or to display close-ups of participants who would be out of the viewer's sight with a single camera. Furthermore, because overlaps between participants captured by multiple conference terminals 100 are eliminated, the number of participants included in the close-up video can be made the same as the actual number of participants being close-up. This improves the visibility of the facial expressions of each participant in close-up.

[0024] Fig. 3 is a diagram showing an example of the functional configuration of a conference terminal 100 according to an embodiment of the present invention. As shown in Fig. 3, the conference terminal 100 includes two camera modules, camera module 10-1 and camera module 10-2 (hereinafter, referred to as "camera modules 10" when not distinguishing between them), an image processing unit 20, an image synthesis processing unit 30, a RAM 70 locally connected to the image synthesis processing unit 30, an audio processing unit 40, a microphone array 41 connected to the audio processing unit 40, an audio output unit 42, an overall processing unit 50, an operation unit 51, a recording processing unit 52, a power management unit 60, a USB device port 80, and a USB host port 81.

[0025] Although FIG. 3 shows an example in which there are two camera modules, the number of camera modules is not limited to a specific number.

[0026] Each camera module 10 captures an image of the conference scene. The camera module 10 includes a lens 11, an imaging unit 12, and a DSP 13.

[0027] The imaging unit 12 is, for example, an image sensor, and generates image data (RAW data) by converting an image collected through the lens 11 into an electrical signal.

[0028] DSP 13 is an abbreviation for Digital Signal Processor, and performs known camera image processing such as Bayer conversion and 3A control on the image data (RAW data) output from the imaging unit 12 to generate image data (YUV data).

[0029] An ultra-wide-angle (fisheye) lens is suitable for the lens 11. An ultra-wide-angle (fisheye) lens allows shooting with as wide an angle of view as possible in the horizontal direction, and allows a single conference terminal 100 to shoot images of many participants.

[0030] In this embodiment, fisheye camera images from two camera modules 10 are joined together to generate a 360° panoramic image.

[0031] The image processing unit 20 performs various types of image processing according to the purpose on the image data (YUV data) output from each camera module 10. In FIG.

[0032] The Affine conversion unit 21 performs a predetermined distortion correction process on the ultra-wide-angle (fisheye) images transferred from each camera module 10. Specifically, the Affine conversion unit 21 loads a correction map stored in advance in the ROM or SSD of the overall processing unit 50 and sequentially performs conversion processing in accordance with the correction map, thereby generating, for each camera module 10, a 180° panoramic image from which distortion specific to fisheye images has been removed.

[0033] Stitch processing unit 22 stitches together two panoramic images by performing stitch processing (joining correction) on a 180° panoramic image generated by Affine transforming the output image from camera module 10-1 and a 180° panoramic image generated by Affine transforming the output image from camera module 10-2. The image generated by stitching becomes a 360° panoramic image (omnidirectional panoramic image).

[0034] The video composition processing unit 30 generates a close-up video and combines the close-up video with a close-up video sent from a subordinate conference terminal 100. To perform these processes, the video composition processing unit 30 includes a person detection unit 31, a face direction detection unit 32, a magnification unit 33, a same person detection unit 34, a generation unit 35, and a composition unit 36.

[0035] The person detection unit 31 detects the presence of a person's (participant's) face in the 360° panoramic image generated by the image processing unit 20, and extracts the area (range) in which the presence of the face (person) is detected. The extraction of the relevant area may generally be performed by detecting the face using a FaceDetect function. However, taking into consideration that participants do not always look at the camera or face forward even during a video conference, the relevant area may also be extracted by detecting the head using a HeadDetect function.

[0036] The face direction detection unit 32 detects the direction in which the faces of people in the area extracted by the person detection unit 31. That is, the face direction detection unit 32 determines whether the participants included in the area are facing forward toward the camera, and if they are facing diagonally, detects the approximate angle.

[0037] The scaling unit 33 performs close-up processing by cropping a specific area (such as an area extracted by the person detection unit 31 or an area specified by the user using the PanTiltZoom function) from the 360° panoramic image generated by the image processing unit 20 and scaling the cropped area to a specified size to match the output format.

[0038] The same person detection unit 34 determines whether or not the same person as the person included in the close-up video generated by the generation unit 35 (described later) is present in the close-up video sent from the lower-level conference terminal 100. Because the upper-level conference terminal 100 and the lower-level conference terminal 100 are used in the same conference room, it is conceivable that the same person may be photographed by both camera modules 10 depending on their arrangement. The presence or absence of the same person is determined to avoid the same person being included twice when the two videos are combined.

[0039] The generation unit 35 generates a close-up video in which some of the videos of a predetermined number of participants are enlarged and arranged based on the video of each participant scaled by the scaling unit 33 and the speech history. That is, the generation unit generates a close-up video including partial videos of a predetermined number of people extracted from the 360° panoramic video, each scaled to a predetermined size.

[0040] FIG. 4 is a diagram illustrating a close-up image. In FIG. 4, (1) shows the positional relationship between the conference terminal 100 (360-degree panoramic camera) and eight people. (2) shows a 360-degree panoramic image p11 captured in the situation of (1) and a close-up image p12 generated by scaling an image (partial image) of some (three) participants in the 360-degree panoramic image p11. The close-up image p12 includes a predetermined number of windows (three in FIG. 4). One window refers to an area that contains a close-up image of one participant. Because the distance between each participant included in the close-up image p12 and the conference terminal 100 is not necessarily the same, the size of each participant in the 360-degree panoramic image p11 may vary. In other words, the image of a participant farther from the conference terminal 100 will be smaller. Therefore, the zoom unit 33 zooms in the partial video including the participant in question so that the size of each participant included in the close-up video p12 is approximately equal (to fit the size of each participant to the window).

[0041] The synthesis unit 36 synthesizes the close-up video generated by the generation unit 35 with the close-up video sent from the lower-level conference terminal 100 to generate a new close-up video. The output video is transferred to the USB device port 80 and then transferred to the higher-level conference terminal 100 or the host device 200 connected via USB in accordance with the UVC protocol. The specific synthesis format and layout will be described later.

[0042] The RAM 70 is used as a dedicated buffer (work memory) when the generating unit 35 and the synthesizing unit 36 perform their respective video processing.

[0043] The audio processing unit 40 performs predetermined audio processing (e.g., codec processing, noise cancellation (NC), etc.) on audio data included in a container of video data received by the host device 200 connected to the conference terminal 100 from a host device 200 at another location. The audio processing unit 40 outputs the processed audio data to the audio output unit 42. The audio processing unit 40 also performs echo cancellation (EC) processing on audio data that is routed back to the microphone array 41 and input while monitoring the audio data to be output to the audio output unit 42. The audio processing unit 40 also performs predetermined audio processing (e.g., codec processing, noise cancellation (NC), etc.) on audio data to be included in a container of video data to be transmitted to a video conference system at another location. The processed audio data is transferred to the host device 200 via USB (UAC protocol). The host device 200 processes the audio data and the video data in a common container. The audio processing unit 40 also uses a beamforming function to identify the direction of sound picked up by the microphone array 41 and generate beamforming information. The audio processing unit 40 transfers the beamforming information to the video synthesis processing unit 30.

[0044] The microphone array 41 includes a microphone array 411 and an A / D converter 412. The microphone array 411 collects the voices of the participants and outputs an audio signal (analog signal). The A / D converter 412 converts the audio signal (analog signal) of the voice output from the microphone array 411 into a digital signal and transfers the converted audio signal (digital signal) to the audio processing unit 40.

[0045] The audio output unit 42 includes a D / A converter 421 and a speaker 422. The D / A converter 421 converts an audio signal (digital signal) transmitted from the host device 200 at the other location into an analog signal. The speaker 422 receives the audio signal (analog signal) converted by the D / A converter 421 and outputs the audio of the participants collected at the other location.

[0046] The overall processing unit 50 performs overall control of the conference terminal 100. The overall processing unit 50 includes a CPU, ROM, RAM, SSD (UFS), etc. For example, the overall processing unit 50 performs mode setting and status management for each module and block according to instructions from an operator.

[0047] The overall processing unit 50 also has functions such as arbitration of system memory (RAM) usage rights and system bus access rights. The overall processing unit 50 also sets the shooting mode of each camera module 10. The shooting mode settings of each camera module 10 may include automatic setting items (e.g., photometric conditions) that are automatically set according to the environment, and manual setting items that are manually set through input by the operator. The overall processing unit 50 also sets the layout processing performed by the video composition processing unit 30. These settings are made by the operator operating the operation unit 51 or via a dedicated app for the conference terminal 100 installed in the host device 200 (sent in the form of commands via USB), and are stored in the memory (RAM) provided in the conference terminal 100.

[0048] The operation unit 51 includes various input devices (for example, a touch panel, operation buttons, a remote control, etc.) The operation unit 51 accepts input of necessary setting information (for example, operation mode setting, start / stop, etc.) by an operator operating the various input devices.

[0049] The recording processing unit 52 has a built-in video CODEC, and encodes the video and audio of the video conference and stores the encoded data in a storage device. The storage device is, for example, a memory device such as an SSD, UFS, or SD card. Such a memory device may be either built-in or external. When using an external device, the conference terminal 100 should be equipped with an SDIO or USB-I / F required for connection. The recorded data generated by the video CODEC is composed of a combination of audio data output from the audio processing unit 40 and video data generated by the video synthesis processing unit 30.

[0050] The power management unit 60 performs power control (power management of USB-VBUS based on Type-C Power Delivery) in accordance with the Type-C Power Delivery standard. The conference terminal 100 can be both a power receiver and a power supplier. Because Type-C ports are widely used, it is unclear what will be connected. That is, it is not always possible to connect a power supply (AC adapter) that meets the requirements, or to connect a conference terminal 100 that can be connected in terms of power supply. Therefore, the main function of the power management unit 60 is to negotiate with a device connected to the Type-C port and determine the power mode. This negotiation procedure is defined in the Type-C Power Delivery standard. All commercially available Type-C power supply devices and devices that operate on Type-C power supply operate in accordance with this standard.

[0051] The power management unit 60 receives power from a lower-level conference terminal 100 connected to the USB host port 81 or from a PD (Power Delivery) compatible AC adapter. The supplied power is used as a driving power source for each module of the conference terminal 100. If there is surplus power in the power supply and a higher-level conference terminal 100 is connected to the USB device port 80, the power management unit 60 supplies power to the higher-level conference terminal 100.

[0052] 5 shows an example of the configuration of the power management unit 60. In FIG. 5, the power management unit 60 includes a power MOS-FET (Field Effect Transistor) 61, a PD controller (Power Delivery Controller) 62, and a power supply unit 63.

[0053] The power MOS-FET 61 functions as a power switch. In this embodiment, power is supplied from the power adapter or the subordinate conference terminal 100, but if the host device 200 has sufficient power supply capacity, power may be supplied from the host device 200 or the superior conference terminal 100 (as with the conventional USB standard). The role of these power supply methods is determined by negotiation between the PD controllers 62, and the direction of power flow is controlled by the switching function of the power MOS-FET 61.

[0054] The PD controller 62 controls the power supply procedure based on the Type-C Power Delivery standard. When devices supporting USB PD are connected, the PD controller 62 performs necessary communication over a dedicated CC line, a function independent of conventional USB data communication. It communicates with the USB PD controller of the connected device (in this case, the upper-level conference terminal, the lower-level conference terminal, or the host device) and makes various decisions (the CC line: a dedicated communication line for the USB PD controller, newly established in USB Type-C). Specifically, the PD controller 62 determines its role—whether it will be a source or a sink—and, if it is a source, determines the power rules, such as the V and A to supply. Once the role and power amount are determined, power transfer begins, and the system is operational, the PD controller 62 monitors overvoltage, overcurrent, and overheating conditions according to its role. If the PD controller 62 detects an abnormality, it implements measures specified in the standard to prevent the problem from spreading.

[0055] The power supply unit 63 supplies approximately 20W to 25W of power to each module of the conference terminal 100. Conventional USB standards require that the USB host device must source power (VBUS) and the USB device device must sink power. However, USB PD eliminates this restriction, allowing power to be supplied from devices with ample power, regardless of whether they are USB hosts or USB devices. Therefore, the power supply unit 63 is configured with a circuit that receives power from an external power source via VBUS as a power sink and distributes the necessary power to each module in the main unit. The power supply unit 63 also functions as a power source and includes a DC-DC converter IC and other components that generate the four voltages required for USB PD: 5V, 9V, 15V, and 20V.

[0056] The USB device port 80 transmits video to an external device using a USB device function. More specifically, the USB device port 80 has a built-in USB device controller and controls protocols such as UVC / UAC to output video data to an external device (the upper-level conference terminal 100 or the host device 200) and input and output audio data. The USB device port 80 is equipped with a Type-C connector to comply with the Power Delivery standard.

[0057] The USB host port 81 receives video from an external device using the USB host function. More specifically, it has a built-in USB host controller and controls protocols such as UVC / UAC to acquire video data from an external device (a lower-level conference terminal 100) and input and output audio data. The USB host port 81 is equipped with a Type-C connector to support the Power Delivery standard.

[0058] 3, the functions of the video processing unit 20, the video synthesis processing unit 30, and the audio processing unit 40 may be realized by dedicated hardware or by software. When realized by software, the functions of the video synthesis processing unit 30 and the audio processing unit 40 are realized by a processing procedure in which a program installed in the conference terminal 100 is executed by a processor (such as a GPU or DSP) corresponding to each of the video processing unit 20, the video synthesis processing unit 30, and the audio processing unit 40.

[0059] Next, a specific situation assumed in this embodiment will be described. Fig. 6 is a diagram for explaining a specific situation in this embodiment. Fig. 6 shows the layout (position) of each participant and each conference terminal 100 at site a. Specifically, there are 12 participants at site a, and the conference terminals 100 are, in order from top to bottom, conference terminal 100-1, conference terminal 100-2, and conference terminal 100-3. The host device 200 is an IWB equipped with a video conferencing application. Each participant is distinguished by an alphabet from A to U.

[0060] The display content of the composite video (the result of combining the videos captured by each conference terminal 100) transferred from site a to site b is a close-up video of the six most recent speakers. Depending on the placement position of the conference terminal 100 and who the most recent speaker is, it is not always possible to get a close-up of all of the six most recent speakers using just one conference terminal 100. Here, the six most recent speakers are assumed to be participants B, C, D, F, S, and T. Furthermore, the previous speaker (before B) is assumed to be participant U.

[0061] FIG. 7 shows an example of close-up images before compositing generated by the generation unit 35 of each conference terminal 100 in this situation. In FIG. 7, image p1 is a close-up image generated solely by the generation unit 35 of the conference terminal 100-1. Image p2 is a close-up image generated solely by the generation unit 35 of the conference terminal 100-2. Image p3 is a close-up image generated by the generation unit 35 of the conference terminal 100-3. As described above, the six most recent speakers are shown in close-up.

[0062] The numerals 1 to 3 following the alphabetic characters assigned to each participant included in the close-up video correspond to the branch numbers of the conference terminals 100 that captured the video of that participant. For example, B1 is the video of participant B captured by conference terminal 100-1, and B2 is the video of participant B captured by conference terminal 100-2.

[0063] The synthesis process of the close-up video by the synthesis unit 36 of each conference terminal 100 is performed cumulatively, starting from the lowest conference terminal 100. In this embodiment, since the conference terminal 100-3 is the lowest, the synthesis unit 36 of the conference terminal 100-3 does not perform the synthesis process. Therefore, the video p3 is transferred as is from the USB device port 80 of the conference terminal 100-3 to the conference terminal 100-2.

[0064] The composition unit 36 of the conference terminal 100-2 composes the video p3 transferred from the conference terminal 100-3 with the video p2 generated in the conference terminal 100-2.

[0065] Fig. 8 is a diagram for explaining an example of the synthesis process in the conference terminal 100-2. Fig. 8 shows an example in which an image p2_3 is generated by synthesis of an image p2 and an image p3 from the conference terminal 100-3.

[0066] The composite result, image p2_3, contains six windows (image areas including participants). On the other hand, the composite target images p2 and p3 each contain a total of 12 windows. The composite process selects six windows from these 12 windows. Specifically, the composite process is performed by selecting either image p2 or image p3 for each participant and connecting the selected images. In FIG. 8, the crosses next to images p2 and p3 indicate that they were not selected by the composite process. The following describes which of images p2 and p3 is selected for each participant. The six most recent speakers (hereinafter referred to as the "most recent six") based on the speech history (information indicating the timing of speech) recorded for each participant in conference terminal 100-2 and conference terminal 100-3 are participants B, C, D, F, S, and T. The most recent six speakers refer to six different people who spoke most recently. For example, if the most recent utterances are "AAABBCCCDEFG," then ABCDEF are the most recent six people. In this case, it is possible to identify that the first AAA, BB, and CCC are the same person using publicly known technology. For example, it is possible to identify the same person by using the direction of the sound (utterance), the identity of the voiceprint, or facial recognition functions.

[0067] Participant B is included in the most recent six participants and is included only in video p2. Therefore, video B2 of participant B included in video p2 is selected as the subject of synthesis. Note that the same person detection unit 34 determines which window the participant in video p2 and video p3 is the same person.

[0068] Participant C is included in the last six participants and overlaps in video p2 and video p3. Therefore, one of the videos is selected using a predetermined selection method. This selection method will be described later. Here, it is assumed that video C2 included in video p2 is selected as the video to be synthesized.

[0069] Participant D is included in the most recent six participants and overlaps in video p2 and video p3. Therefore, one of the videos is selected by a predetermined selection method. Here, it is assumed that video D2 included in video p2 is selected as the subject of synthesis.

[0070] Participant F is included in the most recent six participants and is included only in video p3. Therefore, video F3 of participant F included in video p3 is selected as the subject of composition.

[0071] Participants S and T are included in the last six participants and overlap in video p2 and video p3. Therefore, one of the videos is selected by a predetermined selection method. Here, it is assumed that videos S2 and T2 included in video p2 are selected as the subjects to be synthesized.

[0072] Participant U included in video p3 is not included in the six closest participants. Furthermore, the six participants in the display frame are filled in by the above participants. Therefore, video U3 of participant U is not considered for synthesis. In this way, the most recent speaker is given priority. In other words, the synthesis unit 36 determines the windows (partial videos) to be included in the synthesized close-up video (video p3) based on the timing of the speech of the participants for each window (partial video).

[0073] As a result of the above, the videos B2, C2, D2, S2, T2, and F3 are concatenated to generate the video p2_3 as a composite result. The video p2_3 is transferred from the USB device port 80 of the conference terminal 100-2 to the conference terminal 100-1.

[0074] FIG. 9 is a diagram illustrating an example of a composition process in the conference terminal 100-1. FIG. 9 illustrates an example in which a composition process is performed between a video p1 and a video p2_3 from the conference terminal 100-2 to generate a video p1_2_3. Here, the video p2_3 is a close-up video for the conference terminal 100-1, containing partial videos of a predetermined number of people extracted from videos captured by multiple other conference terminals 100 and resized to a predetermined size. The composition process is the same as that shown in FIG. 8. In FIG. 9, videos B1, C1, B2, S2, T, and F3 are selected to generate a video p1_2_3. The video p1_2_3 is transferred from the USB device port 80 of the conference terminal 100-1 to the host device 200. The host device 200 transmits the video p1_2_3 to the host device 200 at site B. As a result, the host device 200 at site B displays the video p1_2_3.

[0075] The following describes the processing procedures executed by each conference terminal 100. Fig. 10 is a flowchart for explaining an example of the processing procedures executed by the conference terminal 100. The conference terminal 100 that is the focus of attention in Fig. 10 will be referred to as the "target conference terminal 100" hereinafter.

[0076] When the target conference terminal 100 receives a video transfer request from the host device 200 after entering the Ready state, the target conference terminal 100 executes initialization processing (S1).

[0077] Next, the target conference terminal 100 initializes and starts up the camera module 10, the microphone array 41, the audio output unit 42, and their peripheral processing (S2).

[0078] Next, the host device 200 confirms the connection with the other location, initiates the video conference, and begins capturing video using the camera module 10 (S3). At the same time, the host device 200 sends configuration information, such as the operating mode. The configuration information, such as the operating mode, basically specifies the transfer mode (e.g., resolution, frame rate, video data compression method, etc.) according to the UVC (USB Video Calls) standard. This configuration information varies depending on the type of video conference app. Even if the same type of video conference app is used, the configuration information may change depending on the environment and conditions of the video conference. Generally, USB-connected cameras negotiate between the host and camera according to the procedure defined by the UVC standard. Once the host recognizes a USB device, it begins communication with the device according to the protocol defined by UVC. Failure to follow this procedure could result in, for example, video being transferred at an unexpected resolution, causing a mismatch in host operation. Configuration information, such as the operating mode, is transmitted sequentially from the upper conference terminal 100 to the lower conference terminal 100.

[0079] At this time, a 360° panoramic image is generated by the image processing unit 20 based on the image captured by the camera module 10. In addition, the audio processing unit 40 identifies the direction of the sound picked up by the microphone array 41 and generates beamforming information. A certain piece of beamforming information is information indicating a certain direction. Therefore, beamforming information is generated every time a sound is picked up.

[0080] Based on this information (360° panoramic video, beamforming information, etc.), the generation unit 35 generates close-up video (video including partial video in each of six windows) of a predetermined number of people (six people in the case of FIGS. 8 and 9) (S4). That is, the generation unit 35 generates close-up video based on the direction of the sound. More specifically, the generation unit 35 identifies a predetermined number of participants whose most recent speech timing is up to a predetermined order based on the speech timing of each participant identified based on the beforming information (sound direction) as described below, and applies the timing to each partial video of the participants to each window, thereby generating the close-up video. At this time, the generation unit 35 adds the following information (1) to (4) to each window (participant) included in the close-up video.

[0081] (1) Speech history The speech history is information (information indicating the timing of speech in chronological order) that is recorded in RAM 70 in the form of timestamps of past speech histories for each speech direction (i.e., for each participant) using the beamforming function of the audio processing unit 40. The speech history is also used to check for overlapping close-ups of the same participant in the synthesis process. An overlapping close-up refers to the same participant being included in two videos to be synthesized.

[0082] (2) Boundary size in FaceDetect The boundary size in FaceDetect is information recorded in RAM 70 that indicates the size (number of pixels) of the boundary area (rectangular area determined to be the face area when the face is detected) for each participant when the face (or head) position of the participant is detected by the Face (Head) Detect function of the person detection unit 31. The larger this size, the greater the number of pixels in the subject, and so it serves as a measure of image clarity. When overlapping close-ups of the same participant occur during the compositing process, the image with the largest boundary size is selected as the priority.

[0083] (3) Face direction information The face direction information is information obtained by measuring the degree to which the face of each participant is facing the camera (the angle from the front) by the face direction detection unit 32 and the person detection unit 31, and is recorded in RAM 70. When overlapping close-ups occur during the compositing process, priority is given to selecting images that capture the front as much as possible.

[0084] (4) Participant movement history The means for checking for overlapping close-ups of the same participant is not limited to (1), and the participant's movement history may also be used as additional information. The participant's movement history is obtained by accumulating information from the face direction detection unit 32 and storing the participant's past history of changes in face direction in chronological order as additional information. When checking for overlapping close-ups, if the face direction changes in the same direction at the same time, it can be determined that the participants are the same participant. Alternatively, in addition to face direction, information from the person detection unit 31 may be accumulated and the change history of the participant's position information in chronological order may be used as additional information.

[0085] Next, the target conference terminal 100 branches the process depending on whether or not there is a lower conference terminal 100 (S5). The presence or absence of a lower conference terminal 100 can be determined based on whether or not there is a connection to the USB host port 81. If there is a lower conference terminal 100, the process proceeds to step S6; if not, the process proceeds to step S8. In other words, if the target conference terminal 100 is the lowest (end) in the daisy chain connection, the process proceeds to step S8.

[0086] In step S6, the target conference terminal 100 acquires (receives) the close-up video from the lower conference terminal 100 via the USB host port 81. If the target conference terminal 100 is the conference terminal 100-2, it acquires the video p3 in the example of Fig. 8. If the target conference terminal 100 is the conference terminal 100-1, it acquires the video p2_3 in the example of Fig. 9.

[0087] Next, the synthesis unit 36 of the target conference terminal 100 executes synthesis processing on the close-up video generated in the target conference terminal 100 in step S4 and the close-up video acquired from the lower-level conference terminal 100 in step S6 (S7). The outline of the synthesis processing is as explained in FIGS. 8 and 9.

[0088] Specifically, the two close-up videos contain a total of 12 windows, but since only 6 windows can be included in the combined close-up video, it is necessary to narrow down the 12 windows to 6. This narrowing down process is performed as the combining process.

[0089] First, the synthesis unit 36 sorts (arranges) the 12 windows in descending order of most recent speech time based on the speech history of the participants corresponding to each window in each close-up video (S7-1).

[0090] Next, the same person detection unit 34 checks whether the same participant is shown in close-up in the 12 windows (S7-2).

[0091] A facial recognition function may be used to determine whether participants included in two windows are the same person. However, since it is difficult to simultaneously capture frontal video of the same person in both the target conference terminal 100 and the subordinate conference terminal 100, it is possible that applying a facial recognition function alone is insufficient. Therefore, in this embodiment, the same person detection unit 34 compares past speech histories. Since it is unlikely that two people will always speak simultaneously, matching timestamps in the speech histories are likely to indicate that they are the same person. Therefore, the same person detection unit 34 determines that windows with matching timestamps in the speech histories (timing of speech) belong to the same participant. In other words, the same person detection unit 34 determines whether the two close-up videos contain the same participant based on the timing of speech in each window (partial video).

[0092] When the same person is shown in close-up, the composition unit 36 selects one of two windows containing the same person (S7-3), and deletes the non-selected window from the arrangement of windows in sorted order. Details of this selection process (hereinafter referred to as "window selection process") will be described later.

[0093] The synthesis unit 36 concatenates the top six Windows in sort order to generate a synthesized close-up video (S7-4). Note that the synthesis unit 36 adds, to each Window in the synthesized close-up video, the speech history corresponding to that Window, the boundary size in FaceDetect, and face direction information. Note that the speech history is used to detect the same person when synthesizing the close-up video. The speech history is information based on the direction of voice. Therefore, it can be said that the synthesis unit 36 determines the Windows (partial images) to be included in the synthesized close-up video based on the direction of voice.

[0094] Next, the target conference terminal 100 transmits the combined close-up video from the USB device port 80 to the higher-level conference terminal 100 or the host device 200 (S8).

[0095] Steps S2 to S8 are repeatedly executed for each frame image until the conference (video distribution) ends (NO in S9). The repetition cycle is executed in accordance with the frame rate of the camera output image. When an instruction to end the conference is input (YES in S9), the processing procedure in FIG. 10 ends.

[0096] Next, the window selection process described in S7-3 above will be described in detail.

[0097] Fig. 11 is a flowchart for explaining the window selection process. Fig. 11 explains a case where participant A overlaps in two close-up videos. Let An be the window that includes participant A in the close-up video of the target conference terminal 100, and A(n-1) be the window that is detected to overlap with An in the close-up video from the lower conference terminal 100.

[0098] In step S701, the synthesis unit 36 determines whether the boundary size added to An is equal to or larger than a threshold value α.

[0099] If the boundary size added to An is equal to or greater than the threshold value α (YES in S701), the synthesis unit 36 determines whether the boundary size added to A(n-1) is equal to or greater than the threshold value α (S702). If the boundary size added to An is less than the threshold value α (NO in S702), the synthesis unit 36 selects An as the synthesis target (S704). In this case, An has a higher resolution than A(n-1), and it is considered that An makes it easier to read the facial expression of participant A.

[0100] On the other hand, if the boundary size added to A(n-1) is equal to or larger than the threshold α (YES in S702), the synthesis unit 36 compares the degree of frontality of the face based on the face direction information added to An and A(n-1) (S703). When the face direction information is an angle from the front, with the front being 0 degrees, the smaller the absolute value of the face direction information, the higher the degree of frontality of the face.

[0101] If An has a higher degree of frontality of the face (YES in S703), the synthesis unit 36 selects An as the synthesis target (S704), and if not (NO in S703), the synthesis unit 36 selects A(n-1) as the synthesis target (S705). That is, in this case, since both An and A(n-1) satisfy the resolution, the one with a higher degree of frontality is given priority.

[0102] On the other hand, if the boundary size added to An is less than the threshold value α (NO in S701), the synthesis unit 36 determines whether the boundary size added to A(n-1) is equal to or greater than the threshold value α (S706). If the boundary size added to A(n-1) is equal to or greater than the threshold value α (YES in S706), the synthesis unit 36 selects A(n-1) as the synthesis target (S709). In this case, A(n-1) has a higher resolution than An, and it is considered that the facial expression of participant A is easier to read in A(n-1).

[0103] If the boundary size added to A(n-1) is less than the threshold value α (NO in S706), the composition unit 36 compares the boundary sizes added to An and A(n-1) (S707). If the boundary size added to An is larger (YES in S707), the composition unit 36 selects An as the composition target (S708). If not (NO in S703), the composition unit 36 selects A(n-1) as the composition target (S709). That is, in this case, since neither meets the resolution standard, the one with the higher resolution is given priority.

[0104] As described above, according to this embodiment, multiple conference terminals 100 can be arranged near multiple participants, and faces can be displayed enlarged by close-up display, thereby obtaining clear images in which facial expressions are easy to understand. When each conference terminal 100 synthesizes a close-up image, it does not simply connect multiple close-up images, but rather performs synthesis by eliminating overlaps of the same person. As a result, it is possible to prevent the faces of each participant included in the synthesized close-up image from becoming smaller than before synthesis. Therefore, it is possible to improve the visibility of each person's facial expression in an image containing multiple people.

[0105] In this embodiment, the processes executed by the multiple conference terminals 100 connected in series may be used for purposes other than video conferencing.

[0106] Furthermore, the number and placement positions of the conference terminals 100 connected by a daisy chain can be changed depending on the size and layout of the space (such as a conference room) where the conference is held, making it possible to reduce restrictions imposed by the space.

[0107] The functions of the above-described embodiments can be realized by one or more processing circuits. Here, the term "processing circuit" in this specification includes a processor programmed to execute each function by software, such as a processor implemented by an electronic circuit, as well as devices such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), and conventional circuit modules designed to execute each of the above-described functions.

[0108] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to such specific embodiments, and various modifications and variations are possible within the scope of the gist of the present invention as described in the claims.

[0109] For example, aspects of the present invention are as follows. <1> An image compositing system including a plurality of imaging devices, The imaging device is a generating unit that generates a first close-up video including partial videos of a predetermined number of people extracted from the video captured by the imaging device and resized to a predetermined size; a synthesis unit that generates a third close-up video including images of a predetermined number of people based on the first close-up video and a second close-up video including partial images of a predetermined number of people extracted from images captured by another of the imaging devices and resized to a predetermined size; and When the first close-up video and the second close-up video include partial videos of the same person, the combining unit includes one of the partial videos in the third close-up video. A video synthesis system comprising: <2> The second close-up image is an image including partial images of a predetermined number of people extracted from images captured by the other plurality of imaging devices and resized to a predetermined size. Characterized by <1> The video synthesis system described herein. <3> The image captured by the imaging device is a 360° panoramic image. Characterized by <1> or <2> The video synthesis system described herein. <4> The imaging device is a USB device port for transmitting video to an external device using a USB device function; a USB host port that receives video from an external device using a USB host function; characterized in that it has <1> ~ <3> 10. The video compositing system according to claim 1 . <5> The imaging device is A power management unit that controls power supply based on the USB Power Delivery standard, characterized in that it has <1> ~ <4> 10. The video compositing system according to claim 1 . <6> The plurality of imaging devices are connected in series, and an image is transferred from the imaging device at one end to the imaging device at the other end. Characterized by <1> ~ <5> 10. The video compositing system according to claim 1 . <7> the plurality of imaging devices are connected in series, and power is supplied to each of the imaging devices from a single power supply system. Characterized by <1> ~ <6> 10. The video compositing system according to claim 1 . <8> the generation unit generates the first close-up video based on a direction of a sound; the synthesis unit determines the partial video to be included in the third close-up video based on a direction of sound. Characterized by <1> ~ <7> 10. The video compositing system according to claim 1 . <9> The partial image is an image of a range in which the presence of a person is detected among the images captured by the imaging device. Characterized by <1> ~ <8> 10. The video compositing system according to claim 1 . <10> a same person detection unit that determines whether the first close-up video and the second close-up video include partial videos of the same person based on the timing of the person's speech for each partial video; characterized in that it has <1> ~ <9> 10. The video compositing system according to claim 1 . <11> the synthesis unit determines the partial video to be included in the third close-up video based on the timing of speech of the person in each partial video. Characterized by <1> ~ <10> 10. The video compositing system according to claim 1 . <12> when the first close-up video and the second close-up video include partial videos of the same person, the combining unit determines the partial video to be included in the third close-up video based on the size of the partial video before scaling. Characterized by <1> ~ <11> 10. The video compositing system according to claim 1 . <13> When the first close-up video and the second close-up video include partial images of the same person, the combining unit includes in the third close-up video a partial image that is shot in a direction closer to the front of the person. Characterized by <1> ~ <12> 10. The video compositing system according to claim 1 . <14> In a video compositing system including a plurality of imaging devices, The imaging device a generation step of generating a first close-up video including partial videos of a predetermined number of people extracted from the video captured by the imaging device and resized to a predetermined size; a synthesis step of generating a third close-up video including images of a predetermined number of people based on the first close-up video and a second close-up video including partial images of a predetermined number of people extracted from images captured by another of the imaging devices and resized to a predetermined size; Run the combining step, when the first close-up video and the second close-up video contain partial videos of the same person, includes one of the partial videos in the third close-up video; A video compositing method comprising: <15> In a video compositing system including a plurality of imaging devices, a generation step of generating a first close-up video including partial videos of a predetermined number of people extracted from the video captured by the imaging device and resized to a predetermined size; a synthesis step of generating a third close-up video including images of a predetermined number of people based on the first close-up video and a second close-up video including partial images of a predetermined number of people extracted from images captured by another of the imaging devices and resized to a predetermined size; Execute the combining step, when the first close-up video and the second close-up video contain partial videos of the same person, includes one of the partial videos in the third close-up video; A program characterized by: [Explanation of symbols]

[0110] 10-1 Camera module 10-2 Camera module 11 Lens 12 Imaging unit 13 DSP 20 Video Processing Unit 21 Affine conversion section 22 Stitch processing section 30 Video composition processing unit 31 Person detection unit 32 Face direction detection unit 33 Magnification unit 34 Same person detection unit 35 Generation part 36 Synthesis section 40 Audio processing unit 41 Microphone Array 42 Audio output section 50 Overall processing section 51 Operation section 52 Recording processing unit 60 Power management section 61 Power MOS-FET 62 PD controller 63 Power supply section 70 RAM 80 USB device ports 81 USB host ports 100 conference terminals 200 host devices 411 Microphone Array 412 A / D converter 421 D / A converter 422 Speaker [Prior art documents] [Patent documents]

[0111] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-182688

Claims

1. An image compositing system including a plurality of imaging devices, The imaging device is a generating unit that generates a first close-up video including partial videos of a predetermined number of people extracted from the video captured by the imaging device and resized to a predetermined size; a synthesis unit that generates a third close-up video including images of a predetermined number of people based on the first close-up video and a second close-up video including partial images of a predetermined number of people extracted from images captured by another of the imaging devices and resized to a predetermined size; and When the first close-up video and the second close-up video include partial videos of the same person, the combining unit includes either one of the partial videos in the third close-up video. A video synthesis system comprising:

2. the second close-up image is an image including partial images of a predetermined number of people extracted from images captured by the plurality of other image capturing devices and resized to a predetermined size; 2. The video composition system according to claim 1.

3. The image captured by the imaging device is a 360° panoramic image.

2. The video composition system according to claim 1.

4. The imaging device is a USB device port for transmitting video to an external device using a USB device function; a USB host port that receives video from an external device using a USB host function; 2. The video compositing system according to claim 1, further comprising:

5. The imaging device is a power management unit that controls power supply based on the USB Power Delivery standard; 2. The video compositing system according to claim 1, further comprising:

6. The plurality of imaging devices are connected in series, and an image is transferred from the imaging device at one end to the imaging device at the other end.

2. The video composition system according to claim 1.

7. the plurality of imaging devices are connected in series, and power is supplied to each of the imaging devices from a single power supply system; 2. The video composition system according to claim 1.

8. the generation unit generates the first close-up video based on a direction of a sound; the synthesis unit determines the partial video to be included in the third close-up video based on a direction of sound.

2. The video composition system according to claim 1.

9. The partial image is an image of a range in which the presence of a person is detected among the images captured by the imaging device.

2. The video composition system according to claim 1.

10. a same person detection unit that determines whether the first close-up video and the second close-up video include partial videos of the same person based on the timing of the person's speech for each partial video; 2. The video compositing system according to claim 1, further comprising:

11. the synthesis unit determines the partial video to be included in the third close-up video based on the timing of speech of the person in each partial video.

2. The video composition system according to claim 1.

12. when the first close-up video and the second close-up video include partial videos of the same person, the combining unit determines the partial video to be included in the third close-up video based on the size of the partial video before resizing.

2. The video composition system according to claim 1.

13. When the first close-up video and the second close-up video include partial images of the same person, the combining unit includes, in the third close-up video, a partial image captured in a direction closer to the front of the person.

2. The video composition system according to claim 1.

14. In a video compositing system including a plurality of imaging devices, The imaging device a generation step of generating a first close-up video including partial videos of a predetermined number of people extracted from the video captured by the imaging device, each of which is resized to a predetermined size; a synthesis step of generating a third close-up video including images of a predetermined number of people based on the first close-up video and a second close-up video including partial images of a predetermined number of people extracted from images captured by another of the imaging devices and resized to a predetermined size; Run the combining step including, when the first close-up video and the second close-up video include partial videos of the same person, including one of the partial videos in the third close-up video; A video compositing method comprising:

15. In a video compositing system including a plurality of imaging devices, a generation step of generating a first close-up video including partial videos of a predetermined number of people extracted from the video captured by the imaging device, each of which is resized to a predetermined size; a synthesis step of generating a third close-up video including images of a predetermined number of people based on the first close-up video and a second close-up video including partial images of a predetermined number of people extracted from images captured by another of the imaging devices and resized to a predetermined size; Execute the combining step including, when the first close-up video and the second close-up video include partial videos of the same person, including one of the partial videos in the third close-up video; A program characterized by:

Citation Information

Patent Citations

  • Image processing apparatus, image processing method, and image processing program

    JP2012182688A